WI-060 · Platform & Contracts · ranking-replay-calibration
queued P2 low risk Owner: Paul Reviewer: Paul 0% · 0/1 tasks complete
| Code | WI-060 |
|---|---|
| Phase | Platform & Contracts |
| Order | 21 of 93 |
| Story points | 3 |
| Primary surface | Offline replay + policy-comparison tooling (no live calls) |
| Retires | — |
| Depends on | Candidate ranking engine — deterministic features, fixed-point scoring, priority bands, tie-break, explainability, canonical hash |
| Blocks | — |
Let humans evaluate and evolve ranking policies safely: replay recorded trusted results through current/proposed policies, compare versions with explained movements, and gather offline outcome metrics — never auto-mutating production weights.
Replay
Calibration
P4 Deterministic candidate ranking: the model returns evidence/component scores; TroveSnap computes the final ranking.No linked requirements.
Historical replay runs recorded trusted P3 results through current/proposed policy or alternate intent and reports rank/score/band/top-K movement with a reason per movement; policy-comparison report works; outcome/learning-loop metrics collected offline only (publication of any new weight set stays an explicit human step). Per spec §29/§32.13/§35 (tooling rows).
No paid API / infra spend triggered by this item.
queued Sprint: P&C Wave 3: Vision Pipeline & Cost
Edit status / sprint on the ★ Live Board → — changes are logged live with who / when / why.
No human tasks linked.
None recorded yet.
None recorded yet.
None recorded yet.
None recorded yet.
No tool calls recorded.
No files / artifacts recorded.
No log entries yet.