How we measured Jex
Measured 2026-09-28. The comparison vs Jev shown here uses only these STORED numbers -- Jex does not call Jev for anything else (training, priors, new benchmarks), per the founder's decision 2026-10-01. The settings were fixed on development data before the test data was read, and the test was read once.
Systems
- Jex: a trained head on a shared, frozen multilingual encoder (intfloat/multilingual-e5-base), or a fine-tuned model when development data shows it is better for that dataset. The encoder was chosen from four candidates by average score on development splits only.
- Jev (by TypeSafe): called live through its public API at measurement time (2026-09-28), with the label names as the choices. Jev takes no training examples, so its one score is shown in every column. No Jev calls were made after that date for this or any other Jex comparison.
- A zero-shot LLM (openai/gpt-6-luna): asked to pick the label from the label names.
Datasets and rows
| Dataset | Labels | Test rows |
|---|---|---|
| AG News (research licence, pending legal) | 4 | first 400 of the official test split |
| Emotion (dair-ai) (research licence, pending legal) | 6 | first 400 of the official test split |
| GoEmotions (simplified) | 28, multi-label | 400 sampled from test |
| Banking77 | 77 | 400 sampled from test |
| CLINC150 (plus) | 151 | 400 sampled from test |
| TREC (coarse) | 6 | 400 sampled from test |
| SST-2 | 2 | 400 sampled from the validation split (the official test has no labels) |
| Yahoo Answers topics (research licence, pending legal) | 10 | 400 sampled from test |
| MASSIVE intents: English, German, Spanish, Japanese, Chinese | 60 | 400 sampled from each test split |
Test texts that also appear in the training split were removed. The same rows were sent to every system. Samples use a fixed seed (20260928).
Regimes
- 0: label names only. Jex trains on short examples an LLM writes from the label names.
- 8, 32, 128: that many labelled training examples per label, drawn with five seeds; the table shows the first seed.
- Full: the whole training split.
Scoring
Accuracy for single-label sets, micro-F1 for GoEmotions. A cell counts as a win for either side only when the 95% bootstrap interval of the paired difference on the same rows excludes zero; otherwise it is shown as within noise.
What is not covered
- Latency and price are not part of the accuracy table. Jex answers in about 20 ms on CPU.
- Results on your own data will differ. That is why every Jex model reports its own accuracy on held-out examples.