AI StrategistRich Schefren · Strategic Profits

Glossary

Mirror Score

How well does your AI know you is a question with no answer, because nobody can say what a good answer would look like. Move it to decision level and it becomes countable.

Mirror Score (noun). The rate at which a captured version of a person makes the same call that person made, tested on decisions they have already decided and the system cannot see. Measured per domain, not as one figure.

The question, made answerable

Ask anyone whether their AI knows them and you get an impression. It feels like it gets me. It is pretty good at my style. It still does not really understand the business.

Impressions are unreliable here for a specific reason rather than a general one. When you evaluate an output, you supply the missing judgment while you read it. You look at a draft, mentally apply your standard, note the two things you would change, and rate it good. The rating is of you and the machine working together, which is not the thing you were trying to measure.

Moving the question to decision level removes you from the loop.

On calls you already made, that it cannot see:
how often does it make yours?

The answer you gave is fixed. It happened. There is nothing to fill in and no credit to extend, and the comparison is either a match or it is not.

How to run it by hand

This does not need a product. Twenty minutes and a notebook will do.

  1. Pick ten decisions from the last quarter. Real ones, where you remember both the call and the reason. Mix routine with awkward.
  2. Reconstruct each as it stood beforehand. The situation, the options, the constraints. Remove your answer and anything downstream of it.
  3. Ask for the call, and for the reasoning. Both, because a right answer for the wrong reason will not survive contact with the next case.
  4. Record three things. Did the call match. Did the reason match. Did it say it did not know.

The third column is the one people leave out and it carries the most information. A system that produces a confident answer in a domain where it has nothing captured is not scoring badly, it is scoring dishonestly, and the number you get from it is worthless in both directions.

Why one number is the wrong shape

Judgment does not distribute evenly across a business, and neither does capture. A founder might have two decades of pricing calls behind them and have made four hiring decisions ever. One of those domains can be measured. The other has almost nothing to measure.

Average them and you get a figure that is wrong in the specific way that causes damage. High enough to feel trustworthy in general, while concealing the domain where the system is running on nothing at all. Then the system gets pointed at that domain, because the overall number said it was ready.

ReadingWhat it meansWhat to do
High on routine, low on exceptionsIt learned your defaults, not your judgmentCapture the awkward cases. The defaults were never scarce
Low on routine, high on exceptionsAlmost certainly a sampling artifactRerun with more cases before concluding anything
High everywhere, thin coverageIt is guessing well, which is the dangerous kindAudit provenance before extending any autonomy
Right calls, wrong reasonsMatching outcomes without holding the standardTreat as unscored. It will not generalise
Says it does not knowWorking correctly at the edge of coverageNothing. This is the behaviour you want

What a score is not

It is not a percentage of you. A judgment layer at seventy percent on pricing is not seventy percent of a person, and the number does not compound into a claim about how complete the picture is overall.

It is also not a target to optimise. The rate rises naturally as more real calls get captured, and it can be raised artificially by testing on easy cases, which is the sort of thing that happens without anyone deciding to do it. The disagreements are the actual product of the exercise. Every case where the system went one way and you went the other is a ruling that was never captured, sitting there identified, which is the shortest route to a better layer that exists.

Frequently asked

What is a Mirror Score?

A completeness measure for captured judgment. On decisions you have already made, and which the system cannot see the answer to, how often does it make the call you made? That rate is the score. It replaces an unanswerable question about how well the AI knows you with a countable one.

Why measure at decision level rather than by satisfaction?

Because satisfaction measures the output and you supply the missing judgment while you read it. Asked whether the answer is good, you will say yes to work you would have edited. Asked to compare its call with your call on a decision you already made, there is nothing to fill in. The comparison is against a fixed answer you cannot unsee.

Why is a single overall score the wrong shape?

Because judgment is not evenly distributed. Someone can have twenty years of captured pricing calls and almost nothing on hiring. One averaged number hides both facts and produces the worst possible reading: high enough to trust everywhere, and wrong in the places with no coverage. The useful form is per domain, with the empty domains visible. It is also worth keeping separate from articulation density, which measures a task rather than a person and moves on its own.

How do I run this without any special tooling?

Take ten decisions you made in the last quarter, where you remember the call and the reasoning. Give the system the situation as it stood before you decided, with your answer removed. Record whether it matched, and whether its stated reason matched your actual reason. Ten cases is enough to see the shape, and the disagreements teach more than the score does.

What is a good score?

Lower than people expect, and the number matters less than where the failures sit. A high rate on routine calls with a low rate on exceptions is a system that has learned your defaults and none of your judgment, which is the most dangerous configuration, because the defaults are the part that was never scarce.

Where this sits

What is being measured is described at Imprint. Which level of knowing a score actually tests is at the Five Levels of Knowing, and the category is defined at What is Imprinted AI.

Related terms

Where the term comes from

This is one entry in the vocabulary of a longer argument. The full glossary has fifteen terms. The case they belong to runs about 23,000 words, it is free, and there is no email gate on it.

Read The A.I. Business Manifesto

Nothing on this page is for sale. Quote it, argue with it, or pass it on.

Last updated: 28 July 2026