oneshot technical note 1, October 2026
Receipts for one-prompt claims
The oneshot maintainers
Abstract. Reports that a model produced hundreds of results "from one prompt" leave readers unable to tell a single attempt from the best of many, or one prompt from a prompt refined along the way. We describe oneshot, a small library that records every attempt on a task in a hash-chained receipt, commits to the prompt with a salted hash so it can be revealed later, and derives an honest label from the record. On a toy set of exactly checkable problems, the same run reads as "58 of 60 solved" or as "31 of 60 solved in one attempt with one prompt"; both are true, and only the receipt tells them apart.
1. The claim and the missing record
On 6 October 2026 OpenAI published 722 mathematics manuscripts produced by an unreleased model, grouped into 372 families, from roughly 4,000 problems it had posed[1]. A spokesperson said nearly all came from a single prompt given to a single agent, while noting that some may have needed more than one attempt. About 162 manuscripts have a main result checked in a proof assistant. The prompts were not released, the model is not available, and outside researchers asked, in one mathematician's words, for receipts[1][2].
None of this means the results are wrong. It means the sentence "one prompt, one agent" cannot be checked from what was published, and the same sentence fits very different processes.
2. What a receipt records
A run opens with the task, the model, its parameters and a commitment to the prompt: the SHA-256 of a random salt and the prompt. Each attempt adds an event with the hash of its output, whether it passed a check and which check, the tokens, seconds and cost, and the commitment to the prompt it used. A different prompt is a different commitment, so an edit mid-run is visible even while the text stays private. Closing the run adds the totals. Every event includes the hash of the one before, so removing or editing any event breaks the chain from that point.
3. Labels and claims
From the receipt the label follows mechanically: one prompt and one attempt; one prompt, best of N; prompt changed; or no checked result. Over many receipts the library tests four headline claims: every result was a one-shot; no prompt was changed; every shown result was checked; and no tried task was left out.
4. A toy experiment
The library ships a set of small problems with an exact checker (square roots modulo a prime, sums of two squares, digit puzzles) and a pretend model whose chance of being right falls with difficulty. At temperature 0 it gives the same answer every time, so extra attempts cost tokens and add nothing; with sampling, retries help, and a hinted prompt helps more (Table 1, computed in your browser).
| Setting | Shown | One attempt | Attempts |
|---|---|---|---|
| Computing… | |||
5. Publishing without the prompt
A lab may have reasons to keep a prompt private for a while. A redacted receipt still verifies; it shows the commitment instead of the text. When the prompt is revealed with its salt, anyone can confirm it is the prompt that was used, and that it did not change between attempts.
6. Limits
A receipt proves what was recorded, not that everything was recorded; a run should be logged where others can see it, or by a third party. Timestamps come from the recording machine. A check is only as good as the checker: a proof assistant confirms the proof of the statement as written, not that the statement is the intended one[1].
References
- "OpenAI says a secret AI model cracked hundreds of open math problems in one prompt; mathematicians want receipts," Decrypt, 7 Oct 2026. decrypt.co
- "OpenAI says 372 math results came mostly from one prompt," AI Weekly, Oct 2026. aiweekly.co