A Sekiya question · Models & Statistics · Systems & Security

Is the model you deployed actually the model you tested?

Every evaluation makes a quiet promise: that the thing measured is the thing that ships. Sekiya keeps finding that promise broken in interesting places — in release pipelines where identical-looking artifacts produce different decisions, and in evaluation protocols where the test itself manufactures the confidence it reports.

The work under this question asks where the gap opens, how wide it can get, and what a boundary that closes it has to look like.

Objects bearing on it

NoTypeObjectAreaDate ↓State
SK-S-01studyWhen overlapping windows invent predictabilityM&S2026-08-02accepted

Public repositories

← All questions