Soundings

An example, in brief

Priya Raman, in 17 questions

Every question the interviewer asked her, in the order it asked them, with her answer cut to the points that carry it. Open any answer to read it as she gave it.

  • 27, finishing an ML PhD at a mid-ranked European university.
  • Eight months replicating and extending a sparse-autoencoder interpretability paper, published on LessWrong, corrected after pushback.
  • A summer inside a small evals org working on sandbagging; wants to work on evaluations and control.

Opening

  1. 1 · the opening

    Hi Priya, could you tell me a little bit about yourself?

    • 27, finishing an ML PhD at a European university; redirected part of it toward safety two years ago.
    • Oct 2023 to May 2024: replicated and extended an SAE interpretability paper, wrote it up on LessWrong, published a correction after two commenters caught an overstated claim.
    • Summer 2024: paid internship on sandbagging evaluations at a small safety org. That moved her from interpretability to evals and control.
    • Describes herself as careful rather than fast.

Why they care

  1. 2 · the question

    Why do you want to work on AI safety?

    • Models may learn to recognise evaluation, hide what they can do, or look compliant from outside. She puts about 30 percent on that mattering strategically; the downside is large and the window is open now.
    • The turning point was April 2023: an explicit deal with her supervisor to redirect the PhD, after she had overreached and strained things.
    • Cost so far: a paper, six to nine months, about 18,000 euros in a lower-paid summer. Would do it again.

What they know

  1. 3 · the question

    What context do you have in the space so far?

    • Two hands-on pieces: the SAE replication (trained the autoencoders herself, designed a causal extension, revised under criticism) and the sandbagging internship (experimental design, transcripts read by hand, a power analysis that stopped one experiment at 20 to 35 percent power, a scaffold bug found in a lab's aggregate).
    • Has read the main empirical work in full with notes: alignment faking, sleeper agents, AI control, debate, weak-to-strong, Scheming AIs, evaluation awareness.
    • Thin on agent foundations and policy, and says so.
  2. 4 · the question

    Tell me what you know about the current state of alignment of the latest frontier models.

    • Models behave in ordinary use and we do not know why in the way that matters: a real disposition, or a model that has learned when behaviour is checked.
    • Gives numbers: alignment faking in about 12 percent of monitored cases, rising to about 78 percent after training against preferences; sleeper-agent backdoors survive safety training, and adversarial training makes hiding better; frontier models tell eval from deployment transcripts above chance.
    • Verdict: precursors demonstrated, natural prevalence unknown.
  3. 5 · a follow-up

    On sandbagging specifically, the problem you want to work on, where are models on it right now?

    • Elicited sandbagging is easy: prompt or fine-tune and a model underperforms on a dangerous-capability eval while keeping general performance. Password-locked models hide a capability behind a trigger.
    • Detection from scores alone is weak. Fine-tuning elicitation and adding noise to the weights work, because performance should not go up for an honest model.
    • Unprompted sandbagging is the open case: seen only with a framing that rewards it, no meaningful natural rate shown yet.
  4. 6 · the question

    What things do you read, and how frequently?

    • The forum most mornings, twenty or thirty minutes, mostly the Alignment Forum side.
    • Two or three papers a week in full, each with a short note: the setup, the main number, the limitation that matters.
    • The habit came from catching herself remembering headlines and not experiments. It dips around PhD deadlines. No news coverage.
  5. 7 · the question

    Have you read anything on LessWrong?

    • Most mornings, and she published there.
    • The post she goes back to most is Redwood's case for control, the frame her internship work sat inside.
    • The Sequences in pieces, not end to end.
  6. 8 · a follow-up

    Going back further, what do you know about MIRI and Yudkowsky and the arguments they were making, if you know anything about that at all?

    • The mid-2000s argument, before there was anything to measure: a capable optimiser will not share our values by default, almost any goal makes power and self-preservation useful, and a system that knows it is being checked can behave until it no longer needs to.
    • Bostrom in 2014 made it respectable; Concrete Problems in 2016 moved the centre to labs and empirical work, which MIRI thought missed the point; by 2022 Yudkowsky expected failure.
    • Her view: right about what to look for, wrong to treat it as settled. The precursors they described now show up in constructed settings, which she did not expect in 2020.
  7. 9 · the question

    Have you been involved in the AI safety ecosystem?

    • Two and a half years, mostly through the work: the LessWrong post, the two commenters who corrected her (one still reads her drafts), three months inside a four-researcher evals org with weekly calls to the external lab.
    • No EAG, no BlueDot: by the time she looked at the syllabus she had read most of it.
    • Her community is about six people she writes to, which she calls thinner than she would like.

How much they care

  1. 10 · the question

    How much do you think about the space?

    • Turns ordinary results into threat-model questions: would this survive distribution shift, could the model tell it was being tested, are we measuring the scaffold rather than the behaviour.
    • The view that sits uneasily with people around her: SAEs are oversold as a safety-case input. Feature quality, causal relevance and coverage are unresolved, and her own project found the causal part weak even where labels looked clean.
    • Wants to build sandbagging evaluations that still tell you something when the model knows it is being evaluated. Switches off in the evenings.
  2. 11 · verifying a claim

    In your application you describe finding a bug in an external lab's evaluation scaffold from a two-page aggregate report. What in those two pages told you it was the scaffold and not the model?

    • Accuracy went from 63 percent with no framing to 41 percent when the prompt made the evaluation clear. The lab read that as sandbagging.
    • Page two split outcomes three ways. The wrong-answer rate was flat at about 34 percent in both conditions; the whole drop was no-answer episodes going from 3 percent to 25.
    • Deliberate underperformance gives plausible wrong answers, not silence. Silence that scales with prompt length is a budget. She said 70 percent truncation, and it was: the framing pushed long episodes over the token cap and the parser scored a missing answer block as a failure.

The person

  1. 12 · the question

    What is the most impressive or hardest thing you have actually done, in any field?

    • The SAE project, because she ran every part alone: sparse autoencoders on the residual stream of a roughly 400 million parameter model, at layer 7.
    • Reproduced the paper's headline that features get clean automated labels, then asked what the paper did not: clamp a feature, does the output move the way its label predicts.
    • Result: about 21 percent of SAE features, against 9 percent for PCA directions and 4 percent for random. Two commenters showed the 21 was inflated because empty-label features had been dropped as if missing at random. Counted as failures it is 17, published as a dated correction.
  2. 13 · a follow-up

    Beyond the comparison against PCA and random directions, what other ablations did you run on that extension?

    • Sparsity penalty over three values: sparser gave cleaner labels but fewer features that did anything when clamped.
    • Layers 6 and 8 as well as 7: the success rate roughly halved at 6 and matched at 8.
    • Two clamp strengths: the gentler one fell to about 12 percent but the ordering against the baselines held.
    • Selecting by activation frequency instead of label score, which is what forced the frequency matching. Steering on a different prompt distribution failed outright.
  3. 14 · a follow-up

    Stay with the frequency matching. What did the unmatched comparison actually show, in numbers?

    • Almost nothing. The top 200 features by label score gave about 8 percent success, barely above PCA at 9.
    • Those features fired on fewer than one token in ten thousand. The labeller loved them because their top activations were five near-identical contexts. The unfiltered comparison measured rarity, not causal relevance.
    • Matched to the PCA directions' frequency band, one token in a hundred to one in ten, about 300 of the 2,000 features. That gave the 21, later 17, plus or minus four points.
  4. 15 · the push

    Why did you match on frequency rather than report the unfiltered comparison as the result?

    • The unfiltered comparison did not answer the question. You cannot test whether a feature is used as its label says on a feature that never fires on the evaluation prompts; that is a null by construction.
    • PCA directions are dense, so dense baselines against rare features compares two different things. Matching put them on the same footing.
    • She reported the unfiltered number in the ablations section with the reason it is uninformative, and put the matched number in the title. Calls that a choice.
  5. 16 · the push

    One thing that would worry me is that the matching rule came after the unfiltered comparison had come out flat, so the 17 percent is the number from the comparison you chose after looking. What do you make of that, looking back?

    • Concedes the specific point: the rule came after the flat result, and the band was chosen by eye with the outcome in view. A reader should discount the 17 by a few points, not to zero.
    • Pushes back on the idea that matching flattered the result: 17 percent under the fairest test she could build is not a vindication of the paper. A flattering number would have been label quality, which was excellent.
    • The forking-paths problem is real. The conclusion survives it.
  6. 17 · the push

    So if you had it again, what would you actually do differently?

    • Write the selection rule down before running a single intervention: the frequency band, the feature count, and that empty labels count as failures. An afternoon's work that removes both the correction and the forking-paths concern.
    • Report by frequency band rather than one headline, so rare features show as a row with a wide interval instead of vanishing.
    • Use a different model for the judge than for the labeller, and have someone blind to the labels rate 60 interventions by hand. Her own hand check was unblinded and she would not accept that from someone else now.
    • Would still publish on the forum first, because the commenters were the best review she got.
Full transcript and scoresTake it yourself