Module 7 · Lesson 7.3

Building the Team

Who you appoint is the highest-leverage decision you will ever make, because every other decision you make afterward is made by them.

From the Founder

The hardest thing I have had to accept as a leader is that you cannot bring everybody along. I do not mean that cynically. I mean that trying to make everyone happy has, in my experience, produced worse outcomes for everyone than choosing well and telling the truth about the choice. I have kept people in seats they were not right for because letting them go felt unkind, and the unkindness I avoided in the moment got redistributed to their colleagues, to the people we served, and eventually to them, because nobody grows in a role they cannot do. Sometimes you have to take the best course of action, cut the loss, and make the decision that benefits the whole even when it does not benefit some person in the short run. I still do not enjoy that. I have simply stopped confusing my discomfort with their good.

Executive Summary

For an executive, appointments are the highest-leverage decision set he controls, and for a mayor this is not an exaggeration but an arithmetic fact: one person cannot direct 45 agencies, so he directs the people who do. This lesson builds the Cabinet Test — four questions covering capability, candor, complementarity, and climate. It takes selection ethics seriously by steelmanning merit, loyalty, and representation as genuinely competing goods rather than pretending only one is legitimate. It reports the hiring-validity research honestly, including the 2022 reanalysis that lowered nearly every estimate and left structured interviews on top. It extends Edmondson's psychological safety past the slogan, describes the diversity-and-performance evidence accurately including a failed replication, and ends where every mayor ends: a permanent civil service you cannot replace.

Learning Objectives

  • Explain why appointments are the highest-leverage decision set an executive controls, and quantify the claim in the mayoral context
  • Apply the Cabinet Test — capability, candor, complementarity, climate — to a real appointment you are facing
  • Steelman merit, loyalty, and representation as competing goods in selection, and defend a considered ordering of them
  • Report the selection-validity and diversity-performance research accurately, including where the popular claims outrun the evidence

Teaching Manuscript

The Decision That Makes the Other Decisions

Lesson 7.2 left you with a question about your own information: what does the person closest to a problem need to tell you that they will not? If you took the Distortion Trace seriously you found something uncomfortable, and you probably found that the chain ran through one or two specific people.

Now I want to tell you where that chain actually gets set. It is not set in your meetings. It is set the day you decide who sits in the seat. Every meeting after that is downstream.

So let me say the executive point plainly, because most leadership writing buries it in a chapter on talent. For a chief executive, appointments are the highest-leverage decision set you control. Not your strategy. Not your budget. Not your communications. The people you put in the chairs, because every subsequent decision — including the strategy, the budget, and the communications — is made or unmade by them.

For a mayor of New York this is arithmetic rather than philosophy. The Charter makes him chief executive under section 3 and hands him appointment and removal of agency heads under section 6. The Mayor's Office of Management and Budget describes the city as funding roughly seventy agencies with more than 300,000 full-time and full-time-equivalent employees; the Fiscal Year 2024 Mayor's Management Report covers 45 agencies across more than two thousand performance indicators. Now do the honest math on a mayor's calendar. He cannot personally direct 45 agencies. He can barely read 45 agencies. What he can do — the one thing he genuinely, personally does — is choose about that many people and hold them to account. Everything else in the job is a consequence of that choice set.

Which is why the amateur version is so costly. The amateur asks whether he likes the candidate, whether the candidate is impressive in a room, and whether the candidate is loyal. Those questions share a defect: all three are answered by watching how the candidate affects you, when what you need to know is how the candidate will affect four thousand people you will never meet.

So here is the framework, and you should be able to run it from memory in a thirty-minute conversation. Call it the Cabinet Test. Four questions.

One. Can they do the work? Not are they smart, not are they experienced — can they do this work, in this environment, at this scale. Capability is domain-specific and scale-specific, and the most common appointment failure in government and business alike is promoting excellent performance at one magnitude into a job that is a different job at ten times the size.

Two. Will they tell me the truth? Lesson 7.2 established what distortion costs. This is where you buy it down. You are not asking whether the candidate is honest in the abstract; nearly everyone interviews as honest. You are asking for evidence — a specific occasion when this person told a superior something the superior did not want to hear, and what it cost them.

Three. Do they cover what I lack? A team of people who share your strengths shares your blind spots, and blind spots are not additive, they are multiplicative. This is the question ego answers worst.

Four. Can the room they run be honest? This is the one almost nobody asks, and it is the one with the largest downstream effect. You are not hiring an individual performer. You are hiring a climate. Whatever this person does to the willingness of their subordinates to speak will be replicated across every unit they touch, and you will experience it only as strangely late information.

Three Claims That All Have a Case

Before the research, the ethics — because selection is a moral act and the honest version of it involves genuinely competing goods, not one good and two temptations. Three claims press on every appointment. Steelman each before you rank them.

The claim of merit says the seat belongs to whoever can best do the work, and that anything else is theft from the people the seat exists to serve. This is the strongest claim and I hold it first. When a bridge inspector is chosen for a reason other than competence, someone eventually drives over the result. Merit is not a preference; it is what you owe the public. But steelman it properly and you will notice its weak point: merit is not self-evident. Someone decides what counts as qualified, and those definitions carry history. A credential requirement can be a genuine proxy for capability or an inherited filter nobody has examined in thirty years, and telling the difference is real work rather than a slogan.

The claim of loyalty says an executive needs people who will actually execute his program, and that this is not corruption but constitutionalism. Steelman it seriously, because reformers dismiss it too quickly. An elected mayor received a mandate from voters. If his commissioners quietly substitute their own preferences, the voters have been overruled by people they cannot remove. Loyalty in this sense means fidelity to the program you were elected on. The failure mode is that this legitimate claim degrades with terrifying ease into loyalty to the person, which is the mechanism that produces every administration that could not hear bad news — and the mechanism, in the extreme, of every patronage scandal in this city's history.

The claim of representation says a government that governs a diverse city should not be run entirely by people who share one background, and that this is a matter of legitimacy and information, not optics. Steelman it: a cabinet that has never encountered a policy's effects will systematically fail to anticipate them, and citizens who see no one like themselves in a government reasonably wonder whether it is theirs. The failure mode is that representation collapses into arithmetic — a seat filled because of a category rather than because of the judgment the person brings, which insults the appointee and shortchanges the public.

Here is my considered ordering, offered as a position you should test rather than a conclusion you should adopt. Merit is a constraint, not a preference: no one enters the pool who cannot do the work. Representation properly operates on the composition of the pool and on the definition of merit itself, forcing you to search wider and to justify every credential requirement you impose. Loyalty operates last, among the qualified, and it means loyalty to the program rather than to the person — which you can test by asking a candidate what he would do if he concluded your policy was wrong.

Notice what that ordering does. It refuses the false choice that dominates public argument, where one side treats every representation concern as an attack on standards and the other treats every standards concern as a pretext. Both of those are lazy. The serious position holds that the public is owed competence absolutely, and that a leader who searches only where he has always searched has not established that he found the most competent person — he has established that he found the most competent person he looked at.

Now apply it to yourself before we move. Take the last appointment you made. Which of the three claims actually decided it? Not which one you cited afterward — which one moved you. Most leaders discover that loyalty decided it and merit narrated it.

What the Evidence Says About Picking People

There is a real science of selection, and its recent history is a case study in intellectual honesty that every leader should know about, because the field corrected itself in public and downward.

For twenty-five years the reference point was Schmidt and Hunter's 1998 meta-analysis in Psychological Bulletin summarizing eighty-five years of research on selection methods. It was enormously influential and it produced the numbers still quoted in hiring seminars today: work sample tests around .54, cognitive ability tests and structured interviews around .51, unstructured interviews around .38.

In 2022 Sackett, Zhang, Berry and Lievens published a reanalysis in the Journal of Applied Psychology arguing that the standard statistical corrections for range restriction — adjusting for the fact that you only observe job performance among people you actually hired — had been applied in ways that systematically overcorrected. When they redid the estimates with defensible corrections, nearly everything came down. Work samples fell from .54 to .33. Cognitive ability tests fell from .51 to .31. Job knowledge tests came in around .40. Unstructured interviews landed near .19, close to the bottom of the table. And structured employment interviews, at .42, came out at the top of the revised ranking.

Sit with what happened there. A field's flagship finding was revised downward by its own leading researchers on methodological grounds, and the revision changed the ordering, not merely the magnitudes. That is what a healthy discipline looks like, and it is also a warning about confidence in numbers you have not revisited.

The practical instruction that survives all of this is unglamorous and almost nobody follows it. Structure the interview. The same questions, in the same order, for every candidate. Questions anchored in past behavior rather than hypotheticals. A scoring rubric defined before you meet anyone. Interviewers scoring independently before they discuss. That is the whole intervention, and the gap between the structured and unstructured interview in the revised table — .42 against .19 — is the largest free improvement available to most organizations. The reason it is not adopted is not ignorance. It is that senior leaders enjoy unstructured interviews and believe they are unusually good at them.

Then there is what happens after selection, which is where Edmondson's work belongs. Amy Edmondson's 1999 paper in Administrative Science Quarterly defined psychological safety as a shared belief that the team is safe for interpersonal risk-taking, and showed it predicted learning behavior. The finding worth extending is the one that started it: in her earlier hospital research, the nursing units with better performance reported more errors, not fewer. The reporting rate was measuring candor, not quality. Every leader reading a low incident count should feel the floor move slightly.

Notice how badly the concept has been flattened in popular use. Psychological safety is not comfort, agreement, or the absence of hard feedback. Edmondson's own later work is explicit that safety without high standards produces a pleasant, unproductive unit. The construct is about whether speaking up is dangerous, not about whether the work is demanding. A team can be psychologically safe and brutally exacting at the same time — in fact that combination is the target.

Finally, diversity and performance, described honestly, because the popular claims outrun the evidence and a leader who repeats them will be embarrassed by someone who has read the papers. The academic literature has long found that the effects of work-group diversity are mixed and heavily moderated; van Knippenberg and Schippers' review in the Annual Review of Psychology in 2007 is a fair summary of that state of affairs. Informational diversity — genuinely different knowledge and experience — behaves differently from demographic diversity, and both interact with task type and team process. The widely cited McKinsey consulting reports claiming a robust link between executive diversity and financial performance were subjected to a replication attempt by Green and Hand in Econ Journal Watch in 2024 using S&P 500 data, and the authors reported that they could not replicate the findings. Scott Page's argument in The Difference is a formal one about problem-solving under specified conditions, not an empirical claim about firms, and it has been criticized on its own mathematical terms. So state the case you can actually defend: a leader who searches narrowly has not maximized capability, and a cabinet with no experiential range will miss consequences it has never lived. That argument is sound, it is enough, and it does not require a business-performance claim that will not survive scrutiny.

Checkpoint — answer before you read on

Without scrolling back: name the four questions of the Cabinet Test, and say which one most leaders never ask.

Twelve, Seventy, and a Night on a Mountain

Scripture treats selection as a first-order act, and it does something with the sequencing that I have never gotten past.

Luke 6:12-13, NASB 1995: 'It was at this time that He went off to the mountain to pray, and He spent the whole night in prayer to God. And when day came, He called His disciples to Him and chose twelve of them, whom He also named as apostles.'

Take the plain sense before any application. The night of prayer precedes the selection. The text places the most consequential appointment decision in the narrative after the longest recorded period of deliberate seeking. Compare that to how appointments usually get made — late in a full day, under time pressure, from a list assembled by whoever was available.

Then look at whom He chose, because it is the part that unsettles people expecting a competence lesson. Fishermen. A tax collector, which in that setting meant a collaborator with the occupying power. A zealot, from the faction most committed to violent resistance against that same power — placed on the same team. And Judas, whom the Gospels tell us was chosen with full knowledge. This is not a selection manual and I will not pretend it is. What it does is prohibit a certain smugness: the criteria in view were not the ones a search committee would have used, and one appointment ended in betrayal within the narrative's own terms.

Numbers 11 gives the harder and more practical case. Moses, worn down by the whole nation's complaints, tells God he cannot carry these people alone and would rather die than continue. And in 11:16 the instruction comes: 'Gather for Me seventy men from the elders of Israel, whom you know to be the elders of the people and their officers.' Read that qualifier closely. Whom you know to be. Not whom you will now recruit — whom you already know. The seventy were identified by existing, demonstrated standing among their own people. This is selection from observed track record rather than from presentation, and it is the ancient version of the point the validity research keeps making: past behavior in the actual domain beats impression formed in a room.

Then verse 17: 'I will take of the Spirit who is upon you, and will put Him upon them, and they shall bear the burden of the people with you, so that you will not bear it alone.' Theologically this passage has been read many ways and I am not going to overstate it. Organizationally, one thing is unmistakable: the burden is genuinely transferred. They bear it with him. The text does not describe Moses keeping the load and gaining assistants.

And do not miss the moment right after, because it is a psychological-safety text hiding in the Pentateuch. Two men, Eldad and Medad, prophesy in the camp rather than at the tent with the others. Joshua tells Moses to restrain them. Moses answers, in 11:29: 'Are you jealous for my sake? Would that all the LORD's people were prophets, that the LORD would put His Spirit upon them!' The junior man's instinct is to protect the leader's standing by suppressing unauthorized voices. The leader's instinct is to want more of them. That is the disposition this entire lesson is trying to produce in you, and it cannot be faked in a meeting — your team has already worked out which of those two instincts is yours.

Checkpoint — answer before you read on

State what Sackett and colleagues' 2022 reanalysis did to the selection-validity rankings, and the one practical instruction that survives it.

The Cabinet, the Bullpen, and the Workforce You Inherit

Bring it to City Hall, where the structure has features every executive should study.

Start with what a mayor actually controls. Under Charter section 6 he appoints agency heads and may remove an appointee whenever in his judgment the public interest requires it. Under section 7 he appoints one or more deputy mayors with whatever duties he determines — the Charter fixes neither the number nor the portfolios. That is a striking grant. The senior architecture of the executive branch is not prescribed; it is designed by the mayor, which means it is also entirely his responsibility. A first deputy mayor structure concentrates coordination in one trusted person and buys enormous throughput; it also creates a single point through which information flows, and if that person's judgment about what reaches you is poor, you have installed a filter with the authority of a principal. Both are real. Choose knowingly.

Now the history, and the contrast is instructive. John Lindsay took office in 1966 with a young, credentialed, reform-minded team and an explicit agenda of restructuring a sprawling government — the consolidation of dozens of agencies into a small number of superagencies was its signature. The talent was genuine. The complementarity was not: an administration composed largely of people who shared its principals' background and instincts proved poorly matched to the municipal unions, the outer-borough electorate, and the permanent bureaucracy whose cooperation every reform required. A transit strike began the day he was inaugurated. His attempt to expand civilian review of the police was overturned by referendum in November 1966. The transferable lesson is not that Lindsay hired badly. It is that he hired for one dimension.

Michael Bloomberg is the counterexample on selection method, and it is instructive precisely because it is contested. He recruited from business and government on demonstrated executive record, physically restructured City Hall into an open bullpen to shorten information paths, and made his signature appointment after the State Legislature granted mayoral control of the schools in 2002: Joel Klein, an antitrust lawyer and executive with no education-administration credential, who required a waiver from the State Education Commissioner to serve as chancellor. Hold both halves of that. It is a genuine case of merit defined by capability rather than credential — exactly the interrogation of qualification standards this lesson called for. It is also a case of an executive deciding for himself which credentials matter, which is a power that produces Klein and also produces appointments that go badly. Note too what the episode teaches about authority: mayoral control came from Albany and has depended on Albany's renewal ever since. A mayor's most consequential appointment in that domain exists at another government's discretion.

Finally, the constraint no candidate should ever forget. A mayor appoints commissioners. He does not appoint the workforce. New York State's Constitution, in Article V, section 6, requires that civil service appointments be made according to merit and fitness, ascertained so far as practicable by competitive examination. That is a state constitutional command, and it means the overwhelming majority of the roughly 306,000 people the Citizens Budget Commission counted on the city payroll at the end of fiscal 2024 cannot be replaced by a new administration and did not choose you. Only a small fraction of positions sit in the exempt and non-competitive classes.

Understand what that means before you take the oath, because it reframes the entire job. You are not assembling an organization. You are inheriting one — with its own tenure, its own memory, and a reasonable expectation that it will outlast you, since it has outlasted everyone who came before. The civil service system exists for good reasons written in the blood of the patronage era, and it is also genuinely rigid. Your leverage is not replacement. It is the forty-odd people you appoint, the standards you set, the information you make safe to deliver, and the work you push to its rightful owner — which is exactly where Lesson 7.4 begins.

Through the Six Lenses

Evidence levels labeled per the Truth & Intellectual Integrity standard.

Biblical

Interpretation (mainstream reading)

Luke 6:12-13 (NASB 1995) places a whole night of prayer before the selection of the twelve — sequence, not sentiment. The chosen included a tax collector and a zealot on the same team, and Judas. Numbers 11:16 specifies selection from demonstrated standing: 'whom you know to be the elders of the people and their officers.' Verse 17 transfers the burden rather than adding assistants — 'they shall bear the burden of the people with you.' And 11:29 records Moses refusing to suppress Eldad and Medad: 'Would that all the LORD's people were prophets.'

Philosophical

Competing views, steelmanned

Three claims genuinely compete. Merit: the seat belongs to whoever can best do the work — though someone still decides what counts as qualified. Loyalty: an elected executive needs people who will execute the program voters chose, or the electorate is overruled by unremovable staff — though it degrades easily from program to person. Representation: legitimacy and information both suffer when a diverse city is governed from one background — though it collapses into arithmetic when a category fills a seat. Our ordering: merit as a constraint, representation shaping the pool and the definition of merit, loyalty last and only to the program.

Scientific

Consensus with a major published correction

Schmidt and Hunter's 1998 meta-analysis in Psychological Bulletin set the reference numbers for a generation. Sackett, Zhang, Berry and Lievens (Journal of Applied Psychology, 2022) showed range-restriction corrections had systematically overcorrected: work samples fell from .54 to .33, cognitive ability from .51 to .31, and structured interviews at .42 topped the revised table against .19 for unstructured. Edmondson's 1999 ASQ paper defines psychological safety; her earlier hospital work found better units reported more errors — measuring candor, not quality. Diversity-performance effects are mixed and moderated (van Knippenberg & Schippers, 2007); Green and Hand (Econ Journal Watch, 2024) reported failing to replicate McKinsey's results.

Historical

Established record + labeled interpretation

John Lindsay took office in 1966 with a talented, credentialed reform team and consolidated dozens of agencies into superagencies; a transit strike began on inauguration day, and his expanded civilian review of the NYPD was overturned by referendum in November 1966. Cannato's The Ungovernable City is the standard account, and its judgments are interpretation. Michael Bloomberg recruited on executive record, built the City Hall bullpen, and appointed Joel Klein chancellor in 2002 under a State Education Commissioner's waiver after Albany granted mayoral control of the schools.

Influence

Practitioner consensus + research inference

Selection is the highest-leverage influence act available to an executive because it is compounding: each appointee shapes a climate that shapes every person under them. The behavior that most predicts whether a unit tells the truth is what happened to the last person who did. Structured interviewing is also an influence discipline, not merely a measurement one — it constrains the interviewer's own liking bias, which is the mechanism by which charismatic candidates outperform their capability in unstructured settings.

Executive

Established (Charter and constitutional provisions)

Charter section 6 gives the mayor appointment and removal of agency heads; section 7 authorizes deputy mayors with duties the mayor determines and fixes neither number nor portfolio, so the senior architecture is his design and his responsibility. A first deputy mayor structure buys coordination and installs an information filter with principal-level authority. The binding constraint is New York State Constitution Article V, section 6: civil service appointment by merit and fitness, ascertained so far as practicable by competitive examination. Of the 306,248 on-board employees the Citizens Budget Commission counted at the close of FY2024, only a small fraction sit in exempt or non-competitive classes.

Case Study

Bloomberg, 2002: Choosing Capability Over Credential

SITUATION. In 2002 the State Legislature granted the mayor of New York control of the school system, abolishing the Board of Education and giving Michael Bloomberg direct authority over the largest school district in the country. He needed a chancellor. CONSTRAINTS. State regulations set credential requirements for the chancellorship that most business and legal executives could not meet; a waiver from the State Education Commissioner was required. The mayor's authority itself was contingent — mayoral control came from Albany and has depended on legislative renewal ever since. The permanent workforce beneath the chancellor was protected by civil service rules the mayor could not alter. DECISION. Bloomberg appointed Joel Klein, formerly head of the U.S. Justice Department's Antitrust Division and a media executive, with no education-administration credential. Commissioner Richard Mills granted the waiver. ANALYSIS. This is a defensible interrogation of a credential standard — exactly what the merit claim requires when a requirement may be an inherited filter rather than a proxy for capability. It is also the same power that produces bad appointments, because the executive who decides which credentials matter has removed the check that credentials provide. The safeguard is not the rule but the rigor of the process replacing it. DISCUSSION. Which requirement in your own hiring is a real proxy for capability, and which one survives only because no one has asked?

Reflection Questions

  1. Run the Cabinet Test on the person you most recently appointed. Which of the four questions did you not actually ask?
  2. Of merit, loyalty, and representation — which one moved you in your last appointment, and which one did you cite afterward?
  3. Name the person on your team who most often tells you something you did not want to hear. What has that cost them so far?
  4. Which of your requirements for a role is a genuine proxy for capability, and which is an inherited filter you have never examined?

Practical Exercise — The Structured Slate

Take the next role you will fill. Before you meet anyone, write the following on one page and do not revise it afterward: the four to six behaviors that will constitute success in the first year; six behavioral questions, in fixed order, that probe those behaviors and nothing else; a scoring rubric with a written anchor for each score; and the one question that tests candor by asking for a specific occasion when the candidate told a superior an unwelcome truth and what it cost. Have every interviewer score independently before anyone discusses a candidate. Then, after the hire, keep the page. In twelve months, compare your scores to what actually happened and note where your instrument was wrong.

Assessment

1. The fourth question of the Cabinet Test — can the room they run be honest? — is distinctive because it evaluates:
2. Sackett and colleagues' 2022 reanalysis is important primarily because it showed that:
3. The finding that better-performing nursing units reported MORE errors indicates that:
4. The honest position on diversity and organizational performance, as stated in this lesson, is that:
5. New York State Constitution Article V, section 6 constrains a mayor because it:

This Week’s Commitment

Name the one person currently in a seat you would not put them in today. Decide this week which it is — develop, move, or release — write the date, and write the specific conversation you owe them. Then name the next appointment you will make and commit in writing to the structured process you will use before you meet a single candidate.

Identity statement to carry this week: “I am accountable for every person I put in a seat. Their failure is my selection, their success is theirs, and my discomfort about a hard call is not a reason and never was.

Discussion Questions

  • Is loyalty to a program distinguishable in practice from loyalty to a person? Design a question you could ask a candidate that would actually separate them.
  • If a mayor cannot replace the workforce, what are his real levers over 306,000 people? Rank them and defend your first choice.
  • Where should representation operate — on the pool, on the definition of merit, on the final choice, or nowhere? Argue the position you least prefer before stating your own.

Reading List

  • Luke 6:12-16 and Numbers 11:10-30 (NASB 1995) — selection preceded by seeking, and the burden actually transferred
  • Paul Sackett, Charlene Zhang, Christopher Berry & Filip Lievens, 'Revisiting Meta-Analytic Estimates of Validity in Personnel Selection,' Journal of Applied Psychology 107 (2022)
  • Frank Schmidt & John Hunter, 'The Validity and Utility of Selection Methods in Personnel Psychology,' Psychological Bulletin 124 (1998) — read alongside the correction above
  • Amy Edmondson, 'Psychological Safety and Learning Behavior in Work Teams,' Administrative Science Quarterly 44 (1999)
  • Daan van Knippenberg & Michaela Schippers, 'Work Group Diversity,' Annual Review of Psychology (2007)
  • Jeremiah Green & John Hand, 'McKinsey's Diversity Matters/Delivers/Wins Results Revisited,' Econ Journal Watch 21 (2024)
  • Vincent Cannato, The Ungovernable City: John Lindsay and His Struggle to Save New York (2001) — interpretive, and worth reading as such