← Back to Jex

How we measured Jex

Measured 2026-09-28. The comparison vs Jev shown here uses only these STORED numbers -- Jex does not call Jev for anything else (training, priors, new benchmarks), per the founder's decision 2026-10-01. The settings were fixed on development data before the test data was read, and the test was read once.

Systems

Datasets and rows

DatasetLabelsTest rows
AG News (research licence, pending legal)4first 400 of the official test split
Emotion (dair-ai) (research licence, pending legal)6first 400 of the official test split
GoEmotions (simplified)28, multi-label400 sampled from test
Banking7777400 sampled from test
CLINC150 (plus)151400 sampled from test
TREC (coarse)6400 sampled from test
SST-22400 sampled from the validation split (the official test has no labels)
Yahoo Answers topics (research licence, pending legal)10400 sampled from test
MASSIVE intents: English, German, Spanish, Japanese, Chinese60400 sampled from each test split

Test texts that also appear in the training split were removed. The same rows were sent to every system. Samples use a fixed seed (20260928).

Regimes

Scoring

Accuracy for single-label sets, micro-F1 for GoEmotions. A cell counts as a win for either side only when the 95% bootstrap interval of the paired difference on the same rows excludes zero; otherwise it is shown as within noise.

What is not covered