oneshot

$1SHOT on Solana

Show the receipt.

"One prompt, one agent" is now a headline, and nobody can check it. oneshot records every attempt an AI agent makes, keeps the failures on the record, and turns it into a receipt anyone can verify. It's open source and it runs in this page.

FZu9nxS72VUGqCzdezNbPNfbAduMpnoPfY3dUg8pump

A real receipt from the toy solver in this page.

Make a receipt

Give a pretend model 10 to 60 small maths problems with one prompt. Every answer is checked exactly. Then compare what a headline would say with what the receipts say.

Temperature

    One square per problem. Pick one to see its receipt.

    Receipt

    Check a receipt

    Paste any oneshot receipt, or take one from above and tamper with it. Every event carries the hash of the one before, so the check names the exact event that no longer fits.

    Tamper with it:

    Nothing checked yet.

      Publish now, reveal later

      Keep your prompt private and still be held to it. The receipt carries a salted hash of the prompt; when you reveal it, anyone can check it's the same text, word for word.

      1. Write the prompt

      2. Run and publish without it

        The run uses your prompt on five problems. The published receipts show the commitment, never the text.

      3. Reveal, and let anyone check

      What a receipt lets you say

      The label comes from the receipt, not from the press release. Put the badge in your README; it's generated by the same function.

      Paper

      A short note on why one-prompt claims need receipts, and what a toy experiment shows.

      oneshot technical note 1, October 2026

      Receipts for one-prompt claims

      The oneshot maintainers

      Abstract. Reports that a model produced hundreds of results "from one prompt" leave readers unable to tell a single attempt from the best of many, or one prompt from a prompt refined along the way. We describe oneshot, a small library that records every attempt on a task in a hash-chained receipt, commits to the prompt with a salted hash so it can be revealed later, and derives an honest label from the record. On a toy set of exactly checkable problems, the same run reads as "58 of 60 solved" or as "31 of 60 solved in one attempt with one prompt"; both are true, and only the receipt tells them apart.

      1. The claim and the missing record

      On 6 October 2026 OpenAI published 722 mathematics manuscripts produced by an unreleased model, grouped into 372 families, from roughly 4,000 problems it had posed[1]. A spokesperson said nearly all came from a single prompt given to a single agent, while noting that some may have needed more than one attempt. About 162 manuscripts have a main result checked in a proof assistant. The prompts were not released, the model is not available, and outside researchers asked, in one mathematician's words, for receipts[1][2].

      None of this means the results are wrong. It means the sentence "one prompt, one agent" cannot be checked from what was published, and the same sentence fits very different processes.

      2. What a receipt records

      A run opens with the task, the model, its parameters and a commitment to the prompt: the SHA-256 of a random salt and the prompt. Each attempt adds an event with the hash of its output, whether it passed a check and which check, the tokens, seconds and cost, and the commitment to the prompt it used. A different prompt is a different commitment, so an edit mid-run is visible even while the text stays private. Closing the run adds the totals. Every event includes the hash of the one before, so removing or editing any event breaks the chain from that point.

      3. Labels and claims

      From the receipt the label follows mechanically: one prompt and one attempt; one prompt, best of N; prompt changed; or no checked result. Over many receipts the library tests four headline claims: every result was a one-shot; no prompt was changed; every shown result was checked; and no tried task was left out.

      Figure 1. The same 60 problems (seed 7) under five settings. Bars show what each run produced; the headline usually counts the whole bar, the honest one-shot count is the first segment only.

      4. A toy experiment

      The library ships a set of small problems with an exact checker (square roots modulo a prime, sums of two squares, digit puzzles) and a pretend model whose chance of being right falls with difficulty. At temperature 0 it gives the same answer every time, so extra attempts cost tokens and add nothing; with sampling, retries help, and a hinted prompt helps more (Table 1, computed in your browser).

      Table 1. 60 problems, seed 7.
      SettingShownOne attemptAttempts
      Computing…

      5. Publishing without the prompt

      A lab may have reasons to keep a prompt private for a while. A redacted receipt still verifies; it shows the commitment instead of the text. When the prompt is revealed with its salt, anyone can confirm it is the prompt that was used, and that it did not change between attempts.

      6. Limits

      A receipt proves what was recorded, not that everything was recorded; a run should be logged where others can see it, or by a third party. Timestamps come from the recording machine. A check is only as good as the checker: a proof assistant confirms the proof of the statement as written, not that the statement is the intended one[1].

      References

      1. "OpenAI says a secret AI model cracked hundreds of open math problems in one prompt; mathematicians want receipts," Decrypt, 7 Oct 2026. decrypt.co
      2. "OpenAI says 372 math results came mostly from one prompt," AI Weekly, Oct 2026. aiweekly.co

      Docs

      One file, zero dependencies. Node 18+ and the browser. Open any entry.

      createRun(options)start a run
      const run = OneShot.createRun({
        task: "double a number",            // what was asked
        prompt: "Write f so that f(3) == 6.",
        model: "your-model",                // named, or the disclosure score drops
        params: { temperature: 0.7 }
      });
      run.attempt(a)record one try, failures too
      run.attempt({
        output, ok: true,                   // true, false, or null when not checked
        check: "tests",                     // "exact" | "lean" | "tests" | "human"
        tokens: 1500, seconds: 12, cost: 0.02,
        summary: "tried the edge case first",
        prompt                              // only if you changed it; the change is recorded
      });
      run.close({ redact })finish and get the receipt
      const receipt = run.close();               // keeps the last passing attempt
      const hidden  = run.close({ redact: true }); // publish without the prompt
      verify(receipt)is it intact?
      OneShot.verify(receipt)
      // { ok: true, problems: [], id: "9f2c…" }
      // edited: { ok: false, at: 3, problems: ["event 3 was changed, removed or moved"] }
      label(receipt) and badge(receipt)what you may say
      OneShot.label(receipt).text    // "One prompt, best of 3"
      OneShot.badge(receipt)         // an SVG for your README
      disclosure(receipt)what a reader still needs
      OneShot.disclosure(receipt)
      // { score: 6, of: 7, items: [ { key: "prompt", ok: true }, { key: "reasoning", ok: false }, … ] }
      summarize(receipts) and checkClaim(receipts, claim)many runs
      OneShot.summarize(receipts).text
      OneShot.checkClaim(receipts, "one-shot")   // also "one-prompt", "all-checked", "nothing-hidden"
      redact(receipt) and reveal(receipt, prompts, salt)commit now, show later
      const published = OneShot.redact(receipt);
      OneShot.reveal(published, "the prompt", salt)   // { ok: true, receipt } or { ok: false, why }
      replay(receipt, solver)run it again and compare
      OneShot.replay(receipt, (prompt, params, tryNo) => myModel(prompt, params))
      // { ok: true, same: 3, of: 3, different: [] }
      Command linerecord, verify, label, reveal
      cat attempts.jsonl | oneshot record --task "Fix the parser" --prompt "Fix the failing test." --model m > r.json
      oneshot verify r.json
      oneshot label  r.json          # One prompt, best of 2 · checked (tests) · disclosure 6/7
      oneshot redact r.json public.json
      oneshot reveal public.json --prompt "Fix the failing test." --salt <salt>
      oneshot badge  r.json > badge.svg
      oneshot demo --n 40 --attempts 3 --edit-after 2

      Source

      Release 0.1.0, MIT licence. The release itself is a receipt: every file with its size and SHA-256, computed in your browser from the files this page ships.

      Get it

      unzip oneshot-0.1.0.zip

      oneshot-0.1.0.zip the whole repository

      oneshot.js just the library

      README.md the readme

      Changelog

      $1SHOT

      The coin behind the library, live from the market.

      $1SHOT MARKET

      waiting for the market

      Price—

      24h change—

      Market cap—

      24h volume—

      Liquidity—

      Supply1,000,000,000

      Tax0%

      Loading the chart…
      Open the full chart