● Wire · 3 posts · updated 11 OCT 2026Wire · updated 11 OCT RSS ↗
Wire · short notes on AI and LLMs

The wire

Dated notes on models, evals and tooling. Newest first, no hype.

Sunday
Latest

assert(llm) now reports Wilson 95% confidence intervals

Every result card, the live grid and the markdown export now show a Wilson 95% interval next to the pass rate. With 15 questions, 14 correct is not “93.3%”: it is somewhere between roughly 70% and 99%. When a model’s interval overlaps with the leader’s, the card says ≈ statistical tie instead of crowning a winner. Wilson is used over the textbook normal approximation because it stays sensible at small n and near 0% or 100%, which is exactly where small eval sets live.

¶ Permalink ·

Friday

OpenAI Evals goes read-only on 31 October

OpenAI’s hosted Evals turns read-only on 31 October 2026 and shuts down on 30 November 2026. If your regression suite lives there, export the eval definitions and the last good run before the read-only date, while you can still re-run them for a baseline. Keep golden items in plain JSON or CSV with an id, a prompt and an expected answer: that shape moves between tools without a converter. The built-in suites in assert(llm) use exactly that shape; importing your own file is next on the roadmap.

¶ Permalink ·

Tuesday

Two retired Claude Haiku ids now return 404

claude-3-5-haiku-20241022 was retired on 19 February 2026, and claude-3-haiku-20240307 followed on 20 April. A request to either id now fails with 404 not_found_error, so an eval that still lists them scores every row as an error, not a wrong answer. A full row of errors for one model means: check the id before you blame the model.

¶ Permalink ·