Soundings

How it works

A fifteen-minute interview that adapts to you

Spoken or typed. No script. It keeps a belief about you and asks whatever narrows it most. Below, a real run: Priya Raman: PhD, deep context, interpretability detour. For the short version, question by question, Priya in brief.

Priya

Resume

I am finishing a PhD in machine learning at a European university, where my work has focused on representation learning and interpretability. From October 2023 to May 2024, I independently replicated and extended a sparse-autoencoder interpretability paper on a small open model. I published the results on LessWrong and revised the post after commenters identified an overstated claim, including a dated correction from 21% to 17%. In summer 2024, I completed a paid internship at a small AI safety organisation, working on sandbagging evaluations. I conducted a power analysis that led us to stop an underpowered experiment and found a bug in an external lab’s evaluation scaffold from aggregate results. I am now primarily interested in sandbagging, evaluation gaming, and AI control.

the resume goes in

1 · the opening

Hi Priya, could you tell me a little bit about yourself?

Hi. I am Priya. I am 27 and finishing a PhD in machine learning at a European university, on representation learning and interpretability. About two years ago I redirected part of the PhD toward safety work. Between October 2023 and May 2024 I replicated and extended a sparse-autoencoder interpretability paper on a…

2 · a question

Why do you want to work on AI safety?

Because I think the systems we are building may learn to recognise when they are being evaluated, hide what they can do, or pursue goals that look compliant from the outside. I put maybe 30 percent on that mattering in a strategically important way. Not a high number, but the downside is large, and the window is open…

3 · a question

What context do you have in the space so far?

Two pieces of hands-on work and a lot of reading. The SAE replication from October 2023 to May 2024 was on the interpretability side: I trained the autoencoders myself, designed a causal extension, published it, and revised it under criticism. The internship in summer 2024 was on evaluations: experimental design for…

4 · a question

Tell me what you know about the current state of alignment of the latest frontier models.

My honest summary: the latest models are well behaved in ordinary use and we do not know why in the way that matters. Post-training rewards the behaviour we want, and the open question is whether it produces the disposition or a model that has learned when the behaviour is checked. The evidence for the second is real…

5 · a follow-up

On sandbagging specifically, the problem you want to work on, where are models on it right now?

Split it into three. Elicited sandbagging is easy: prompt or fine-tune a model to underperform on a dangerous-capability eval and it will, keeping general performance, and password-locked models show a capability can be hidden behind a trigger and recovered with a little fine-tuning. Detection from scores alone is…

6 · a question

What things do you read, and how frequently?

The forum most mornings, twenty or thirty minutes with coffee, mostly the Alignment Forum side and whatever is being argued about that week. Two or three papers a week in full, with a short note afterwards: the setup, the main number, and the limitation I think matters. That habit came from catching myself remembering…

7 · a question

Have you read anything on LessWrong?

Yes, most mornings, and I published there. The post I go back to most is Redwood's case for control, the argument that you can get safety from protocols that assume the model is adversarial, because it is the frame my internship work sat inside. I have read a fair amount of the older material too, the Sequences in…

8 · a follow-up

Going back further, what do you know about MIRI and Yudkowsky and the arguments they were making, if you know anything about that at all?

Roughly this. Yudkowsky and what became MIRI argued from the mid 2000s, before there was anything to measure, that a capable optimiser would not share our values by default, that almost any goal makes power and self-preservation useful, and that a system which knew it was being checked could behave until it no longer…

9 · a question

Have you been involved in the AI safety ecosystem?

Modestly, for about two and a half years, and mostly through the work rather than events. The LessWrong post was my first real contact; the two commenters who corrected me were the first people in the space I argued with properly, and one of them still reads my drafts. The internship put me inside a small evals org…

10 · a question

How much do you think about the space?

A lot, in a particular way. Mostly I turn ordinary results into threat-model questions: would this survive distribution shift, could the model tell it was being tested, are we measuring the scaffold rather than the behaviour. That last one is the habit the internship left me with. The thing I have come to think that…

11 · verifying a claim

In your application you describe finding a bug in an external lab's evaluation scaffold from a two-page aggregate report. What in those two pages told you it was the scaffold and not the model?

The shape of the drop. The report had accuracy by condition: 63 percent when the model was told nothing, 41 percent when the prompt made it clear the run was an evaluation. That looked like sandbagging and the lab read it that way. But the second page split outcomes into correct, wrong, and no final answer, and the…

12 · a question

What is the most impressive or hardest thing you have actually done, in any field?

The SAE project, because I ran every part of it alone and it has the most experiments. From October 2023 to May 2024 I trained sparse autoencoders on the residual stream of a small open model, about 400 million parameters, at layer 7, reproduced the paper's headline that the features get clean automated labels, and…

13 · a follow-up

Beyond the comparison against PCA and random directions, what other ablations did you run on that extension?

Four that mattered. I swept the sparsity penalty over three values; the sparser dictionary gave cleaner labels but fewer features that did anything when clamped, so label quality and causal use pulled in opposite directions. I ran the intervention at layers 6 and 8 as well as 7; at 6 the success rate roughly halved,…

14 · a follow-up

Stay with the frequency matching. What did the unmatched comparison actually show, in numbers?

Almost nothing, which was the surprise. Clamp the top 200 SAE features by label score and success was about 8 percent, barely above PCA at 9. Those features fired on fewer than one token in ten thousand, and clamping something that rare barely moves an output on any prompt you can test; the labeller loved them because…

15 · the push

Why did you match on frequency rather than report the unfiltered comparison as the result?

Because the unfiltered comparison was not answering the question. The question was whether a labelled SAE feature is used by the model in the way the label says, and you cannot test that on a feature that never fires on the prompts you evaluate on; you get a null by construction, not a finding about the model. PCA…

16 · the push

One thing that would worry me is that the matching rule came after the unfiltered comparison had come out flat, so the 17 percent is the number from the comparison you chose after looking. What do you make of that, looking back?

That is fair on the specific point. The rule was chosen after I had seen the flat result, and the band I matched on was chosen by me, by eye, with the outcome in view. Any reader is entitled to discount the 17 for that, and I would put the honest discount at a few points rather than zero, because the confound's…

17 · the push

So if you had it again, what would you actually do differently?

Three things, in order of cost to me. First, write the selection rule down before running a single intervention: the frequency band, the feature count, and that empty labels count as failures. That one change removes both the correction and the thing you just raised, and it would have cost me an afternoon. Second,…

From the resume alone

AI safety context

0%

0 of 7 asked

Readingprior
AI safety contextprior

Mission alignment

0%

0 of 2 asked

Mission alignmentprior

Research judgment

0%

0 of 6 asked

Research judgmentprior

Traits

0%

0 of 10 asked

Judgmentprior
Bias resistanceprior
Agencyprior
Growthprior
Ambitionprior
Interpersonalprior
Integrityprior
Technical skillprior

each bar is where the truth could still sit, 1 to 10; the mark is the best guess

Read an exampleTake it yourself