Powered by RiskAICenter · Non-Profit · Advancing Agentic AI in Finance

Free Tool · No Login Required

Evaluate any black-box AI system

Works for any AI — a chatbot, a classifier, a trading agent, a recommendation engine — not just investment strategies. Enter what you fed the system and what it produced, in order. Text and numbers are both fine, on either side. Everything computes in your browser, per Aldridge's zero-instrumentation, covariance-based methodology — nothing is ever uploaded.

Black-box inputs & outputs

One row per observation, in the order you tested them. Left = what you gave the system. Right = what it produced. Either side can be text or a number. At least 6 complete rows are required. Have a whole file of input/output pairs instead? Use bulk CSV evaluation.

Input
Output
Cov · Earlier Half
–
Input–output covariance, first subsample
Cov · Later Half
–
Input–output covariance, second subsample
Δ Covariance
–
Absolute change, later minus earlier
Δ %
–
Relative change vs. earlier half
Verdict
–
Per "Regret Equals Covariance"

Covariance: earlier half vs. later half

If the bars are close in height, the input–output relationship held constant. A large gap means it shifted.

Interpretation

How each row was scored

Text is embedded into a shared vector space (a signed, character-aware hashing embedding) and reduced to a scalar with a fixed random projection; numbers embed as themselves. This is what covariance is actually computed on.
#InputTypeEmbedded ScalarOutputTypeEmbedded ScalarSubsample

Methodology: every input and output is embedded into the same fixed-dimensional vector space, entirely in your browser — numeric entries embed as themselves, and text is embedded with a deterministic, signed hashing-trick embedding that includes character-level subword features (fastText-style), then L2-normalized. Each embedding is reduced to one scalar with a fixed random projection, the standard way to collapse a high-dimensional embedding while preserving its relative structure. The full ordered sample is then split in half; the sample covariance between input and output scalars is computed separately in each half; and the two covariances are compared. Per Aldridge's "Regret Equals Covariance" framing, a stable covariance across the two halves indicates consistent, low-regret system behavior, while a shifting covariance is read as behavioral drift — rising regret. This is a reference implementation of the general public framing of that methodology, intended for teaching and quick triage rather than as a substitute for a full model-risk review.

Need this running continuously on a live system?

The free evaluator scores a hand-entered sample, once. The Pro and Team plans connect to your pipeline and re-run this covariance-stability test on every new batch of inputs and outputs automatically — with drift alerts.

See Pro & Team plans →