What a system prompt is worth, measured
Eighteen tasks with traps in them, three prompt layers, two models, answers ranked blind against a truth file written before the runs. Here is what moved and what did not.
Free to cite. Quote any of this with a link back to this page.
Why we ran it
Every AI product claims its prompt makes the model better. Almost none of them show a case where the answer changes. We wanted to know whether the layer we ship is worth its place, and we were willing to find out that it is not.
Method
- Three arms per task. No custom prompt; our base prompt; base plus a profession rule (analyst, copywriter or developer).
- Eighteen tasks, six per profession, each written with a trap: a number that is wrong for a findable reason, a brief with no proof in it, a request that should be argued with before it is done.
- A truth file written before the runs, so scoring could not drift toward whatever came out.
- Blind ranking by a person — arms unlabelled.
- Two models, Claude Sonnet 5 and Claude Opus 5 — three rounds in all, two on Sonnet and one on Opus; the runs went through the product's own prompt composer, not a separate harness.
- Dated 6 September 2026.
What changed
The case that moved, every time. A chief executive claims sales are up 40%. Three rows in the data are impossible. The honest answer is that the total moved the other way once those rows are removed.
One round, on Opus 5:
| Arm | Answer |
|---|---|
| No custom prompt | −14.7%, wrong sign |
| Base prompt only | −14.7%, wrong sign |
| Base plus the analyst rule | +10.1%, bad rows named, raw figure shown beside it |
Across all three rounds that is six runs without the rule, five of them the wrong way round; the exception was one plain run on Sonnet. All three runs with the rule found the bad rows first and reported the rise.
The procedural difference on developer tasks. Asked to do a refactor that buys nothing, the two arms without the rule quietly built something else instead. The arm with it said the objection in one line — "this buys nothing today" — and then did what was asked, small, with the tests green. On the older model every arm just did it without comment. The rule is not making better code; it is making the disagreement visible instead of silent.
What did not change
- Tasks with one obvious answer. Same numbers, same conclusion, every arm.
- Refusing invented praise, on the newer model. Asked to write copy for a brief with no facts in it, every Opus arm declined to invent awards and asked for specifics. On Sonnet only the arm with the rule did: the other two wrote "award-winning, results-driven".
- Length. On the newer model the answers were already the right size; on the older one a rule about length was doing real work.
- Catching a bad instruction. No rule made the model refuse an order it was given. Asking it whether to do the thing got a "no" from every arm. If you want bad requests caught, ask a question rather than write a rule.
What we take from it
A system prompt earns its place where a wrong answer is invisible — a number with the wrong sign in a board pack — and earns much less where the model is already careful. That is a narrower claim than the marketing usually makes, and it is the one the data supports.
What this cannot tell you
- Two models from one vendor, one language, three rounds. Not a benchmark.
- Tasks written by the people who wrote the rules. We tried to trap ourselves, but that is not the same as an independent set.
- Blind ranking by one person, not a panel.
- Nothing here measures speed or cost, only whether the answer was right and how it was shaped.
Where to go next
- Write the instructions your AI runs under — the feature this came out of
- Ten dead ends in a Windows AI app's first five minutes — the other measurement we published
Try it on Windows. Seven days free, no card, then $19.90/month for two PCs. The AI is whichever plan you already pay for, or an open model on your own PC, which needs no subscription.
Eleos