Skip to content
CCAF Preparation

Why your CCAR-F practice score doesn't predict your real score

The exam reports a scaled score on a 100–1,000 range and Anthropic has not published the scaling function. What that means for every practice score you will see, including ours.

Every practice platform for this exam reports a number that looks like the real one. Ours does too. It is worth understanding exactly what that number is, because the honest answer changes how you should use it.

What the real exam reports

You pass CCAR-F at a scaled score of 720 on a 100–1,000 range. Your report also gives percent-correct by domain.

Scaled scoring is not the same as percentage correct. It exists so that scores are comparable across different sittings that drew different items — if one form happens to include harder questions, the scaling compensates so that the same ability produces the same score. Getting 42 of 60 right does not map to a fixed number, because it depends on which 42.

Anthropic has not published the function that does this mapping. That is normal for a proctored certification; publishing it would make the exam easier to game.

The consequence

No third party can reproduce the real score. Not this site, not the paid question banks, not the courses. Anyone showing you a number on a 100–1,000 scale is applying their own approximation and presenting it in the exam's clothing.

Ours is a linear map from percentage correct onto the 100–1,000 range, weighted by domain in the blueprint proportions. It is a defensible approximation and it is not the real thing. A 690 here does not mean you would fall 30 points short. It does not mean you would pass either.

There is a second problem underneath the first. A practice score is only as good as the items it is computed from, and the real item bank is confidential. Our questions are written against the same public task statements, and twelve of them are the official samples reproduced verbatim — but a bank written from a blueprint will always have a slightly different difficulty distribution from the bank the blueprint was written for. Two approximations, stacked.

What the number is actually good for

Three things, all of them comparative rather than absolute.

Finding your weakest domain. This is the real product of a mock exam. Domain weights are uneven — 27% for Agentic Architecture versus 15% for Context Management — so a domain you are weak in costs you differently depending on which one it is. The per-domain breakdown tells you where the next hour of study belongs, and that recommendation is robust even when the headline number is not.

Measuring movement. The gap between your first and third mock, on the same bank, is meaningful even if neither number maps to reality. Improvement is a relative measurement, and relative measurements survive a shaky scale.

Rehearsing the format. Sixty items in 120 minutes is two minutes per question, on scenario-based items that take real reading. Knowing what that pace feels like — and having already discovered which of the four scenarios you slow down on — is worth doing at least twice before the real thing.

What it is not good for

Deciding whether to book the exam. A practice score cannot tell you that, and treating it as a gate produces both failure modes: people who delay a booking they were ready for, and people who book on the strength of a number that was never measuring what they thought.

Better signals for readiness: have you built the four things the exam guide recommends building? Can you explain, without looking, why a guardrail belongs in the orchestration layer rather than the prompt? When you get a practice question wrong, is it because you did not know the fact, or because you picked a workaround over a mechanism? The last one is the real test — the exam is a judgement test, and judgement does not show up cleanly in a score.

How to read ours

The mock exam here draws 60 questions from a bank of 318 in the official blueprint proportions and reports a scaled figure against the real 720 threshold. We show it because a bare percentage makes the domain weighting invisible, and the weighting is the thing that matters.

Read the per-domain table underneath it first. Then look at the headline number, and hold it loosely.