Profile
Back to NewsBack
Hacker News 8 min
Reader Mode
What if AI worked at 1.000.000 tokens per seconds?

What if AI worked at 1.000.000 tokens per seconds?

13 hours ago

One Million Tokens a Second

The original animated film, with narration and music. Tap the picture to play. This is a thought experiment, not a measurement of any current model. The reading version below adds two interactive calculators.

Prefer to read & explore? ↓
01 / WHAT THE NUMBER MEANS

“A million a second” can mean four things.

CONTEXT CAPACITYHow much fits
The text a model can hold in one request, including its reply. A size, not a speed.
INPUT PROCESSINGHow fast it reads
Prompt tokens can be processed largely in parallel. Huge inputs still take time, and the input rate is different from the output rate.
AGGREGATE THROUGHPUTHow much a system serves
10,000 streams × 100 tokens/sec = 1,000,000 tokens/sec in total. Toy arithmetic: it says nothing about when any one stream finishes.
SINGLE-AGENT OUTPUT · OUR DIALHow fast one agent writes
One agent writing a million tokens a second. In standard generation, each token depends on the ones before it.

Anthropic’s Claude Opus 5.5 overview, for example, lists a 1M-token context window, a much smaller output cap per request and “moderate” comparative latency. Memory size is not writing speed, and it does not mean a million-token reply fits in one request.

Engineering helps, within limits. Services reach big totals by batching many requests, as the 2023 vLLM paper describes, and speculative decoding can accept several drafted tokens in one pass. Neither removes a genuine chain in which step two needs the result of step one.

A token is a chunk of text, often part of a word, so token counts are not word counts.

TRY IT / THE IMAGINED DIAL

Turn the dial.

Pick a workload, then slide the imagined speed from 100 to 1,000,000 tokens per second.

1,000,000 tokens/sec
1001K10K100K1M
OUTPUT TOKENS1,000,0001,000 story endings × 1,000 tokens each
WRITING ONLY1 secondat 1,000,000 tokens/sec
Writing timeLog scale
1 sec1 min1 hour1 day

At 1,000,000 tokens/sec, 1,000 story endings (1,000,000 output tokens) take about 1 second to write.

All four budgets at both ends of the dial +
Writing-only time at illustrative speeds, not measurements.
WorkloadOutput tokens100 tokens/sec1,000,000 tokens/sec
1,000 story endings × 1,0001,000,0002 h 46 min 40 s1 second
40 app drafts × 20,000800,0002 h 13 min 20 s0.8 seconds
10,000 critiqued candidates × 3003,000,0008 h 20 min3 seconds
10,000 rehearsals × 1,00010,000,00027 h 46 min 40 s10 seconds

Imagined comparisons, not measurements of any model. Times count generated output only: no reading, ranking, tool calls, permissions, tests, deployment, people or experiments. A budget is a writing allowance, not a finished product.

Notice

Parallel streams can accelerate these independent jobs too. But each draft still needs time on its own stream: matching aggregate throughput does not match completion time. One fast agent matters most when each step waits on the last.

02 / FOUR THOUGHT EXPERIMENTS

Cheap drafts move the hard part.

Each budget counts written output only. None is a finished product, a verified discovery or a real result.

A map of possible endings

1,000 × 1,000 = 1,000,000 tokens · 1 second of writing

Ask for an ending to your story and get a thousand: hopeful, dark, strange. Nobody reads a thousand, so the useful version maps them for you to explore, then blends the two you like.

Still scarce: taste. Only you know which ending is yours.

Software for a neighborhood tool library

40 × 20,000 = 800,000 tokens · 0.8 seconds of writing

Describe a lending app for shared tools and forty drafts exist before you finish the sentence. Screens could adapt, from borrowing a ladder to a repair-day sign-up. Tests still run on their own clock, and if none checks the seven-day due date, all forty can pass while lending ladders for seventy.

Still scarce: a clear spec, and tests of what “working” means.

Ten real experiments from ten thousand ideas

10,000 × 300 = 3,000,000 tokens · 3 seconds of writing

Hunting for a better catalyst? An agent could propose and critique ten thousand candidates. Then everything waits at the bench: reactions may run for hours, cultures for days, field trials for a season. Ideas from one model can also share one blind spot.

Still scarce: physical evidence, and choosing which ten experiments earn lab time.

Rehearsing a hard conversation

10,000 × 1,000 = 10,000,000 tokens · 10 seconds of writing

Before talking to your landlord about the lease, an agent could play the conversation ten thousand ways: stubborn landlord, generous landlord, you when tired. Use it like a flight simulator, for practice and blind spots. Ten thousand rehearsals of the wrong person are a confident mistake, not a prophecy.

Still scarce: fidelity to the real person, and your own practice.

03 / THE SERIAL BOTTLENECK

10,000× faster writing is not 10,000× faster work.

Take the forty app drafts: 800,000 output tokens. Compare an illustrative 100 tokens per second with the imagined million, then add one fixed check after writing that speed does not touch, such as a test run.

60 seconds
01 min1 hour1 day
Illustrative 100 tokens/sec2 h 14 min 20 s

Writing 99.3% · Checking 0.7%

Imagined 1,000,000 tokens/sec60.8 seconds

Writing 1.3% · Checking 98.7%

WritingFixed checkEach bar shows where its own total goes.
WRITING SPEEDUP10,000×1,000,000 ÷ 100 tokens/sec
END-TO-END SPEEDUP132.57×8,060 ÷ 60.8 seconds

With 60 seconds of checking, the batch takes 2 h 14 min 20 s at 100 tokens/sec and 60.8 seconds at 1,000,000 tokens/sec. That is 132.57× faster end to end, and checking is 98.7% of the faster total.

The same 800,000 tokens with four check times +
Total time = writing + one fixed check, counted once.
Fixed check100 tokens/sec1,000,000 tokens/secEnd to end
None2 h 13 min 20 s0.8 seconds10,000×
60 seconds2 h 14 min 20 s60.8 seconds132.57×
15 min2 h 28 min 20 s15 min 1 s9.88×
24 h26 h 13 min 20 s24 h 1 s1.09×

Toy model: total time = output tokens ÷ speed + one fixed check, counted once for the whole batch (not per draft) after writing ends. It does not simulate parallel tests, cost, energy or draft quality.

With a one-minute check, the job drops from 8,060 to 60.8 seconds: about 132.57 times faster, not 10,000. At zero the full 10,000× returns; at a day the gain nearly vanishes. Whatever you do not speed up becomes nearly all the remaining time, the logic of Amdahl’s law. Real requests have more such steps: OpenAI’s latency guide notes that very large prompts, tool calls and network trips add delays of their own.

04 / WHAT GETS PRECIOUS

When generation gets cheap, judgment gets precious.

Conceptual drawing, not data: generation widens the options; judgment decides what passes.

Every experiment above ends in the same place. What stays scarce is judgment, wearing different hats:

  • TasteWhich option is yours.
  • Specs and testsWhat “working” means.
  • PracticeSpeed cannot learn it for you.
  • Physical evidenceThe world answers at its own pace.
  • FidelityA simulation is only as good as its model.
  • Cost and energyIs this worth running at all?

Fast does not mean cheap: every token runs on hardware someone pays to power. More options do not guarantee a good one either. Drafts from one model and one set of assumptions can be correlated, or all wrong together. Volume measures output, not understanding.

05 / USE IT THIS WEEK

Before you ask for more, decide how you will choose.

No imaginary dial required. Next time you hand work to an AI agent:

  1. Write down what “done” means.One testable sentence, like “ladders are due back in seven days,” not “handles loans.”
  2. Set criteria before reading options.Two or three, so the most fluent draft does not win by default.
  3. Find the slow step.Name the check that speed will not touch. Shorten or automate it where possible, without skipping the validation each result needs.
  4. Ask for disagreement, not volume.Request options built on different assumptions, then ask what would make all of them wrong.

If your slow step is deciding what to measure or which bets to make, that is strategy more than tooling. Private consulting helps clarify the larger system and where your next move matters ↗

Sources, dates and boundaries.

This Field Note adapts ideas from echohive’s film One Million Tokens a Second; it is an edited companion, not a transcript. The speed is imagined, and no source below measures, claims or predicts it. They support only the distinctions between capacity, reading, serving and writing.

Open the four sources +
  1. Anthropic, Claude Opus 5.5 overview. Official documentation. Lists a 1M-token context window, a separate maximum output per request and “moderate” comparative latency. It does not describe a million output tokens per second.
  2. Anthropic, Context windows. Official documentation. Describes the context window as working memory that includes the generated response: a capacity, not a speed.
  3. OpenAI, Latency optimization. Official API guide. Output generation commonly dominates latency; very large prompts still matter, and tool or network calls add their own delays.
  4. Agrawal et al., Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. arXiv, 2024. A historical foundation, not a current benchmark: parallel prompt processing (prefill), token-by-token decoding and batching.

The 2023 vLLM paper linked in section 01 is also a historical foundation. All workloads, speeds, budgets and the bottleneck model are illustrative arithmetic, not benchmarks, forecasts or product specifications. Sources checked October 2, 2026.

KEEP THE CURIOSITY. STRENGTHEN THE PRACTICE.

Understand the shift.
Practice the judgment.

Learn at your own pace, think it through with others on Sundays, or get help with your own direction. The Architect membership includes the full Get Amplified collection plus the 1000x Lab: mostly live discussion with replays, not a coding class. Private consulting is separate and not included.

LEARN AT YOUR OWN PACE

Get Amplified

Move from interesting ideas to a repeatable practice: AI workflows, research methods, markets and attention.

Explore the field guide ↗
THINK IT THROUGH TOGETHER

1000x Lab

Bring your questions to the Sunday conversation. Explore changing ideas and their implications through live discussion and replays.

See how the Lab works ↗
APPLY IT TO YOUR OWN DIRECTION

Private consulting

Step back from the tools. Clarify the larger system, what deserves your attention, and where your next move has leverage.

Explore private consulting ↗

Ready to make this a practice? Compare the Get Amplified and 1000x Lab options on Patreon, where the current tier details are explained.

Explore membership on Patreon ↗
Chat with me