I am so glad I finally published that note yesterday. Probably had it on my brain because of this coding assessment I was going to take.
Anyway, my timing is amazing.
I just took the assessment and completely blanked.
I got a 0 on it.
No kidding.
An actual zero.
Not a disappointing score.
Not a score below the cutoff.
ZERO.
Out of 600.
Which, now that I'm theoretically over the immediate horror of seeing and living it, presents a methodological problem. The 0 is real. It is actual data. Under those conditions, on that assessment, I produced zero scored performance.
In a totally non-defensive way, I might question what we're entitled to infer from that observation, though, given what was observable immediately before it. Both literally, and in yesterday's research note.
I'm not trying to prove that I secretly deserve a 500/600, although if anyone knows anyone at Anthropic (ouch!) who could somehow manage to forgetfully lose a digital score, I won't complain. Instead, I've decided to ask this much more interesting measurement question that was prompted by an exceptionally inconvenient data point:
Sarah V.
So, what else can we observe?
In the days before the assessment, I was practicing Python. A. lot. Not to pretend I'm a software engineer. Not to produce beautiful code. Mostly to get SQL out of my head for the day. I practiced the kinds of basic programming problems I expected to ecounter (because that's what the email said): loops, lists, dictionaries, classes, retrieving values, counting things, filtering things, figuring out whether I needed a count or a list or a total before I started typing. Essentially Python kindergarten.
And I was getting better. But more importantly, immediately before I started the assessment, I could do things that I apparently could not do once the assessment started. It's a fine line, but it matters.
If all someone had was the assessment (and, given the score, that's probably all they will have), they would have a perfectly legitimate observation: I received a 0/600. Again, I demonstrated no scored performance on this assessment.
What they wouldn't know, is that shortly before taking it, and while trying to psych myself up to actually begin, I could reason through many of the component skills the assessment was presumably designed to test. They would not know what I understood but failed to retrieve in a timed test (in which I wasn't allowed to make use of my substantial stack overflow skills, btw). They wouldn't know where the breakdown happened (besides to the user in front of the keyboard, obviously). They would not know whether I never actually acquired or implemented the relevant knowledge, if I misunderstood the problems, if I couldn't translate what I knew (my logic skills are amazing) into code, if I ran out of time, if I froze under the assessment conditions, or if I experienced some particularly disastrous combination of all of those.
My score won't tell them that. Not that it claims to. Which is where this gets sort of interesting (and by interesting I mean interesting on top of absolutely horrifying).
Again, the measurement isn't false because it's incomplete. I really did earn that amazing score. The assessment truly, successfully captured something that happened under a particular set of conditions. But the minute someone uses that observation to make a claim about what I can do, instead of what I did under those conditions, they've made an inference. (yay stats!)
Now I want to know whether there is enough information to support that.
This is weirdly familiar. Yesterday's note was about the technical system I built for my dissertation research. I spent tons of time getting a fairly complicated sequence of tools to work together. I had participants moving through Prolific and Qualtrics, interacting with the OpenAI API through FastAPI, generating transcripts and snapshots, and eventually producing structured data that could survive all that and still know who it belonged to.(And again, I'm a social scientist.)
So that worked. But then I had to confront a much less satisfying question: was the thing I had successfully built actually giving me the evidence I needed to answer the question I cared about?
Apparently, I just needed roughly 24 hours to implement this as a real-world, participant-facing problem using myself.
Cool.
To be clear, I am not arguing that any of this means the assessment is bad. Maybe Anthropic needs all the world's non-machine learning engineers to be able to look at an unfamiliar coding problem, figure it out, and implement a solution correctly under time pressure without outside help (beyond the documentation that is decidedly less helpful when one freezes on a timed test.) If that's the capability they need to identify, then my spectacular 0 is actually extremely useful information.
But that's also a much more specific claim than "Sarah can't code." Conveniently, I'm making this the distinction I'm interested in.
This is something we do all the time. We observe an outcome and use it as evidence about the thing that produced it. It makes sense. Usually, we have to. We can't observe everything that happened inside someone's head (see previous notes about changing my methodology). We can't reconstruct every decision, false start, retrieval failure, moment of confusion, change in representation, or sudden realization that happened along the way. We measure something we can see.
A score of 0 is quite clear to see.
An answer, a final artifact, a completed task, and maybe a self-report about what happened can all be fairly clear to see. So we use them, and then we infer backwards.
That is often completely reasonable. Often, the endpoint is exactly what we care about. But sometimes the process is the phenomenon. Where have I heard that? Somewhat inconveniently to this assessment, this is where I've spent the last several months ending up with my dissertation.
I've been interested in what happens when people work with genAI on problems that aren't fully specified at the beginning. I'm not looking at only whether AI helps them produce a better final answer, but whether interacting with it changes what they think the problem is, what possibilities they notice, what they abandon, what they return to, and where they eventually end up.
The problem is that a final artifact is a terrible observer for most of that. Two people can arrive at similar answers through completely different trajectories. They can produce very different answers afterf incredibly similar answers. Their final response might look like it hasn't changed even though their understanding of the problem shifted substantially along the way. Or, something might look like a dramatic change at the end even though it emerged through a series of small, locally sensible moves. (Oh man, just wait until I publish my "why Pinky from Pinky and the Brain is so recognizable to me" research note!)
Endpoints are completely real. They just aren't the whole event. And apparently neither is 0/600 (and I am so glad I have my interest in process already documented before writing this).
What I had immediately before the assessment wasn't the "true score." I don't know what my true score is, and after this assessment I'll maybe not accept guesses yet. What I had was additional evidence. There was a trajectory before the measurement.
I started with some programming knowledge and a lot of experience thinking logically about data, but relatively little recent practice writing /fixing Python syntax. I made extremely basic mistakes. I confused things, I wrote things in the wrong order, and I assume I caused Python some emotional distress. But then I corrected things.
Patterns that initially required a lot of thought became easier to recognize and I got better at identifying what kind of structure a problem required when tables apparently don't exist, before I tried to write it. I could explain why I was getting answers wrong andretrieving things I wasn't able to retrieve earlier. Immediately before the assessment, I was making mistakes, but I was also demonstrably doing things that I hadn't been able to do as reliably a few days ago.
Then the assessment happened. And whatever happened during that time (which at least someone will know because they recorded audio, video, AND my screen), resulted in: 0.
If I retain only the endpoint, I lose all that movement. Again, this doesn't make the endpoint wrong. It just changes the kinds of questions I can answer with it. And that's starting to feel like the methodological problem underneath a surprising amount of what I care about.
I've been using the phrase "representational trajectories" in my dissertation to describe how people's understanding of a problem changes through interaction. I'm beginning to think the word trajectory is earning its keep.
It's not just: Did the representation change? It's: What kinds of movement occurred? When? In response to what? What persisted? What disappeared and came back/ Where did something that looks obvious in the final answer actually enter the process? And what becomes invisible when we only preserve the thing at the end?
A score of 0 is seriously good at making this particular problem visible because it is hilariously stark. There isn't ambiguity in the presentation. It's very clearly ZERO. OUT OF 600. Which reasonably looks like the end of an argument.
In one sense, it probably is. I took an assessment as part of an application process. The organization now has a data point it can use to make a selection decision. I would strongly prefer that particular data point had been literally any positive number instead. But the decision problem belongs to them, and shockingly, I'm not part of the methodological process. There are perfectly good reasons to care about exactly what the assessment captured.
I tend to propose a different flavor of research question, though. I want to know what happens to thinking while it's happening. And for that question, an endpoint can't tell me enough.
Yesterday I wrote about spending months building a system that worked and then realizing that technical success didn't mean I had captured the construct I actually cared about.
Today, I produced the most technically unsuccessful outcome available to me short of not figuring out how to click the link to the assessment, and realized essentially the same thing from the other side.
The data is real. They're allowed to decide whether it's the data they need for the decision they're making. I get to remain extremely interested in what is lost between the the trajectory and the score.
And sure, I'm also allowed to hope someone in my network knows someone at Anthropic and passes this along. I told you I was approaching this in a totally non-defensive way. I never said I stopped wanting the fellowship.
Anyway, if Anthropic calls, none of this happened.
Sarah, remember: this is allowed to be unfinished. If it were finished, it wouldn’t be a Research Note.
As always, these notes reflect my own intellectual work. AI was used for organization and conceptual scaffolding, both as support of and as an object of iterative inquiry into the creative process and cognition.