Observation
The current score response is the rounded probability weighted expected bucket, as specified in ADR-003 and enforced by engine tests. A validation-only experiment found that this contract and exact bucket accuracy measure different outcomes.
On the frozen tickets validation split (150 rows), using only the filtered train split for fitting:
| Model and budget |
Rounded expected bucket |
Probability argmax |
Train-selected majority baseline |
| bge-small-en-v1.5 INT8, 8 shots per class, minCalibration 8 |
37.33% |
49.33% |
58.67% |
| bge-small-en-v1.5 INT8, 32 shots per class, minCalibration 8 |
55.33% |
68.00% |
58.67% |
For the 32-shot arm, rounded output was bucket 1 on 109 of 150 validation rows, while the truth contained 44 bucket 1 rows. This is consistent with averaging probability mass from the endpoint buckets into the middle bucket. No test labels were consulted and no package behavior was changed.
Decision to investigate
Preserve the existing score contract. Evaluate ordinal MAE and probabilistic calibration alongside exact bucket accuracy. Compare an ordered head or an explicitly named top bucket output that callers can opt into. Do not silently change the meaning of score or promote a model using the same split used to choose its policy.
Acceptance
- Define the target metric for ordinal use cases and a separate categorical bucket metric.
- Fit candidate policies using train and calibration only; select on frozen validation and check transfer before scoring test once.
- Maintain backward compatibility for
score, with tests for both semantics if an opt-in output is added.
- Record exact-match, MAE, ECE, latency, and per-bucket confusion counts with pinned model hashes and split hash.
Observation
The current
scoreresponse is the rounded probability weighted expected bucket, as specified in ADR-003 and enforced by engine tests. A validation-only experiment found that this contract and exact bucket accuracy measure different outcomes.On the frozen tickets validation split (150 rows), using only the filtered train split for fitting:
For the 32-shot arm, rounded output was bucket 1 on 109 of 150 validation rows, while the truth contained 44 bucket 1 rows. This is consistent with averaging probability mass from the endpoint buckets into the middle bucket. No test labels were consulted and no package behavior was changed.
Decision to investigate
Preserve the existing
scorecontract. Evaluate ordinal MAE and probabilistic calibration alongside exact bucket accuracy. Compare an ordered head or an explicitly named top bucket output that callers can opt into. Do not silently change the meaning ofscoreor promote a model using the same split used to choose its policy.Acceptance
score, with tests for both semantics if an opt-in output is added.