Free Tool · No Login Required
Bulk-evaluate a whole file of inputs and outputs
Upload a CSV with one column of what you gave an AI system and one column of what it produced — a chat log, a Q&A export, a batch of model calls. Pick a window width K, and a K-row window slides across the whole file, recomputing input–output covariance at every position so you can see where (not just whether) the relationship held or drifted. Parsing and every calculation happen in your browser — the file is never uploaded anywhere.
1. Upload a CSV
Any CSV with a column of inputs and a column of outputs works. Row order matters — rows are read top to bottom, in file order.
Need an example? Download a sample CSV — 100 rows of a retail-investor Q&A log (Question / Answer columns), the same shape this tool expects.
Set your acceptable covariance range
Rolling covariance —
Interpretation
Window-by-window detail
| Window | Rows | Covariance |
|---|
Methodology: every input and output in the uploaded column is embedded into the same fixed-dimensional vector space and reduced to a scalar with a fixed projection (the same pipeline as the single-sample evaluator — see /tools/ai-evaluator for the full writeup). A window of K consecutive rows is then slid one row at a time across the whole file, and the ordinary sample covariance between the input and output scalars is computed inside each window, producing one covariance estimate per window position. Per Aldridge's "Regret Equals Covariance" framing, a rolling covariance that stays close to its own average is read as consistent, low-regret system behavior, while a covariance that trends or swings widely is read as behavioral drift. Rather than relying on a single fixed threshold (or on sign changes alone, which can flag a harmless flip near zero as a "change"), the stability verdict above is driven by the acceptable-covariance range you set with the sliders: a window only counts against stability if its covariance actually falls outside the range you defined. The default range (mean ± 1 standard deviation of the rolling series) is a starting point, not a recommendation — narrow it for a stricter read or widen it for a looser one. This is a reference implementation of the general public framing of that methodology, intended for teaching and quick triage rather than as a substitute for a full model-risk review.
Need this running continuously on a live system?
The free evaluator scores a file you upload by hand. The Pro and Team plans connect to your pipeline and re-run this rolling-window test automatically as new data arrives — with drift alerts.
See Pro & Team plans →