10 Comments
User's avatar
Pawel Jozefiak's avatar

The jump from 18.5% to 76% on needle-in-haystack is the stat that matters most here. Numbers are nice, but what changed for practical work?

I loaded my entire blog archive into Sonnet 4.6's 1M window last week. Not a synthetic benchmark - actual retrieval across years of messy, overlapping content. It pulled references I'd forgotten writing, correctly cross-linked themes, zero hallucinated sources. That felt different from anything I'd tried before.

The DRAM cost angle you cover is the part most reviewers skip entirely. 1M context doesn't come free. Someone's paying for that memory, and the 172% price surge explains why inference costs aren't dropping as fast as model capabilities improve.

Full test results here: https://thoughts.jock.pl/p/sonnet-46-two-experiments-one-got-personal

Nibmeister's avatar

I've been experimenting with Opus 4.5 and 4.6 for translating the Septuagint from the Greek, chapter by chapter. With 4.5, after I finished the translation of Job, I asked it to compile the entire book from the chapter translations and translation notes, and it broke after getting about half of the translation notes compiled. A second prompt and it was able to finish compiling the translation notes.

For Opus 4.6, I did the same process with the book of Esther, and encountered the same problem, so I must suspect that Gab, where I did this translation, must not have enabled the larger context window. There was no control on Gab for this, nor could I turn on the reasoning mode on Gab. I will try again on openrouter.

J Scott's avatar

"is increasingly about retrieval quality at scale"

This is what I want. Fidelity makes it most useful to me.

SKY DOG's avatar

I have 128GB of DDR5 6000MHz RAM that is not timing-compatible with my workstation. I should put it in a safe deposit box.

Gridhunter's avatar

Quite bragging about your retirement plan, SkyDog!

SKY DOG's avatar

(insert meme of noble looking down at the plebs from the balcony)

Nym Coy's avatar

Oh hey, now that you mention it, I ran my whole novel through it and it actually commented on the middle for once.

Vox Day's avatar

I just want to know when we'll actually get our promised ONE MILLION TOKEN context window operational in Claude Opus 4.6.

The Gray Man's avatar

Yeah Google has had it solid with 64k output for two years

Jordamøn's avatar

"Bigger is not automatically better"

We already have it, but Opus 4.5 was more efficient with its 200k tokens that 4.6 has been so far with its million.