assert(llm) now reports Wilson 95% confidence intervals
Every result card, the live grid and the markdown export now show a Wilson 95% interval next to the pass rate. With 15 questions, 14 correct is not “93.3%”: it is somewhere between roughly 70% and 99%. When a model’s interval overlaps with the leader’s, the card says ≈ statistical tie instead of crowning a winner. Wilson is used over the textbook normal approximation because it stays sensible at small n and near 0% or 100%, which is exactly where small eval sets live.