scienthoon/jev-ood-calibration
7 0PythonMITUpdated 5 days ago
Independent calibration test of TypeSafe's Jev on a task it cannot have seen: 900 rule-generated support tickets (choice / score / boolean) plus 3 public benchmarks via Vercel AI Gateway. Raw responses, ECE with noise floor, temperature refit, per-type sign of miscalibration. Reproducible for ~$0.06.
Our verdict: worth borrowing ideas from
Contains scripts and methodology for OOD calibration testing of Jev, offering reusable evaluation pipelines and ECE calculations that can be adapted to his own models
Filed under agent harnesses in our directory.