Soundings

The rubric

How the interview runs and what it tests

One 20-minute voice call. Nothing is read from a list: every line is written from what the person just said. A separate grader places them on each trait afterwards.

The 20 minutes

  • What they want the interview to highlightMinute 1
  • IntroMinutes 1 to 3
  • Their story and probes into their work: where the time went, what failed, the hardest arguments against the things they didMinutes 3 to 11
  • AI safety context: what they read, what it showed, the state of the fieldMinutes 11 to 16
  • Mission: their timelines, what they would work on, why that, what they do if rejectedMinutes 16 to 19
  • CloseMinute 20

As measured on Akshay's own call and the four latest simulated ones: the intro was over between minute 1.7 and 3.8, the work had until minute 11 to 11.6, context until 15.8 to 16.5. Eleven to seventeen questions in all.

How a follow-up is made

The moment they stop, one model writes the next line with the whole transcript in view. A second model works in the background after every answer: what changed in its picture of the person, what is worth pushing on, which kinds of question come next. It hands over material, never a line. The writer follows these principles; each example is from Akshay's own interviews, with his verdict.

  • Never ask what is already known

    “How long ago would you say you started getting into AI safety?”

    Bad: her intro had said early 2023. “You literally already know.”

  • Every question is the next step in their story, and tests a trait without saying so

    “And how did you get intros to them?”

    Trying to see how he thinks and solves problems. A high signal question.

  • Model the best they could have done in that time, then ask about the gap

    “Why did you lose 6 weeks? Why did you not catch that earlier? What were you doing during those 6 weeks?”

    “Much harder-hitting.”

  • A premise they can correct, instead of asking if they care

    “And this was because you had expertise in those areas already?”

    A sly probe to see if he will say no, actually, because the mission was so compelling. He did not.

  • Fewer follow-ups, the best ones

    “If frontier labs adopted this check, what larger safety decision would you want it to make more reliable?”

    Bad: “too deep in the follow-up. This just seems unnecessary.”

What is tested

Each trait is a ladder from 2 to 10. Hover over the question mark beside one, or press it, to see what each level means and what it looks like.

Their path and main work

What they actually did, how rigorous it was against how hard it was, and where the time went.

Their story in order, one wide probe on the main work told to go long, one dig into a thing they named, then one push on a decision they made.

Traits read here

  • Research judgment
  • Technical skill
  • Judgment
  • Agency
  • Integrity
  • Growth
  • Bias resistance

AI safety context

What they have read and taken from it, how today's models fail, the arguments, the people and the orgs.

A cluster of about four minutes: what they read, what a work showed, what they built on it, the state of the field.

Traits read here

  • AI safety context
  • Reading

Mission alignment

How much they care about the world beyond themselves, and how much that care shapes what they work on.

Read off the reasons for each move in their story, then asked directly: their timelines, what they would work on, why that, and what they do if rejected.

Traits read here

  • Mission alignment
  • Ambition
See it workTake it yourself