Trainable text classifiers

The classifier you own.

Train it on your data in minutes. More accurate than the generalists, answers in under 20 ms of server time (measured p50 12.7 ms) -- in over 50 languages. Your examples never train anyone else's model.

No GPU needed2 to 200 labelsMulti-labelDelete any time
Classifier "support-tickets" · 4 labels

“Ich komme nicht mehr in mein Konto, das Passwort funktioniert nicht.”

account_access
99.9%
bug
0.1%
billing
0.0%
Trained on 120 English examples in 2.8 s. Answered in 13 ms.

What you get

Seconds

To train

Jex trains a small model on top of a shared multilingual encoder. 10,000 examples train in about a minute on CPU.

Under 20 ms

Per classification (server time)

Measured p50 12.7 ms on the serving CPU. One encoder pass and a tiny head. Send up to 50 texts per call; large batches under heavy concurrent load queue behind each other -- see the docs' limits.

50+

Languages

Train in English, classify German, Spanish, Japanese or Chinese. Or train in any of them.

Report card

For every model

Accuracy, per-label precision and recall, a confusion matrix and a confidence cut, all on examples the model never saw.

Private

Your data stays yours

Examples are used for your classifier only and discarded after training unless you ask us to keep them. Export or delete at any time.

API or upload

No-code or code

Paste a CSV on this page, or POST examples from your app. The same classifier either way.

How to use it

  1. Get an API key

    During the preview, the demo below gives you a key that lasts a day. Account keys arrive with launch.

  2. Send labelled examples

    A list of {text, label}, or {text, labels} for multi-label. Add a group (a conversation or document id) so near-copies never land on both sides of the test.

    curl -X POST https://getjex.dev/v1/classifiers \
      -H "Authorization: Bearer $JEX_KEY" -H "Content-Type: application/json" \
      -d '{"name": "tickets", "examples": [
            {"text": "I was charged twice", "label": "billing"},
            {"text": "Export crashes on Safari", "label": "bug"}]}'
    # 202 {"id": "c_…", "status": "queued"}
  3. Read the report card

    GET /v1/classifiers/{id} returns the status and, once ready, the report: accuracy and macro-F1 on a held-out test split, per-label precision and recall, the confusion matrix, and the confidence at which it is right 90% of the time.

  4. Classify

    curl -X POST https://getjex.dev/v1/classifiers/c_…/classify \
      -H "Authorization: Bearer $JEX_KEY" -H "Content-Type: application/json" \
      -d '{"texts": ["Refund my annual plan"]}'
    # {"results": [{"label": "billing", "confidence": 0.94, "confident": true, "scores": {…}}]}
  5. Improve it

    The report lists the test examples it got most wrong. Add examples for the weak labels and train again; each training returns a fresh report so you can compare.

  6. Delete it

    DELETE /v1/classifiers/{id} removes the model, its report and any kept examples, and returns a deletion receipt.

Or: classify instantly, with no training examples

Send just your label names (and, better, a one-line description of each). POST /v1/classify answers immediately -- every call, including the very first one on a brand-new label set. The first few calls on a label set you have never sent before come back as mode: "zero_shot_warming" (an instant answer from label-embedding similarity, no training yet); once Jex has generated and fit a small model for that exact label set in the background (seconds, cached from then on), every later call is mode: "instant_head" and more confident. Common label shapes (spam/ham, sentiment, support intents, topics) are pre-warmed, so they never show the cold-start mode at all.

curl -X POST https://getjex.dev/v1/classify \
  -H "Authorization: Bearer $JEX_KEY" -H "Content-Type: application/json" \
  -d '{"labels": {"billing": "a payment, invoice, or charge issue",
                "bug": "something in the product is broken",
                "feature_request": "a request for a new capability"},
       "texts": ["I was charged twice this month for one seat"]}'
# first call on this label set (zero-shot, warming a model in the background):
# {"results": [{"label": "billing", "confidence": 0.4092,
#     "scores": {"billing": 0.4092, "bug": 0.3302, "feature_request": 0.2605}}],
#  "mode": "zero_shot_warming", "ms": 51.0}
# a later call, same label set, once it has warmed:
# {"results": [{"label": "billing", "confidence": 0.804,
#     "scores": {"billing": 0.804, "bug": 0.1375, "feature_request": 0.0585}}],
#  "mode": "instant_head", "ms": 11.8}

2-200 labels, 25+ word-or-two label descriptions recommended for the best cold-start answer. No training examples are billed or stored; generation happens on our side. See the docs for errors, limits and the trained-classifier mode above.

Try it

Examples, one per line: label, text

The demo allows 2 classifiers of up to 2,000 examples, kept for 1 day. The report needs about 25 examples per label.

Report card

Train the classifier to see accuracy and macro-F1 on a held-out test split, the most-common-label baseline, and precision and recall per label.

Classify a text

How Jex compares

Jex vs Jev (by TypeSafe) on 13 public datasets, STORED numbers from a 2026-09-28 run -- no new calls to Jev. Every system saw the same 400 rows from each official test split. Jex's settings were chosen on separate development data before the test was read.

Jex ahead (95% interval excludes zero)Within noiseJev ahead
No examples58.2 vs 76.0Jex avg vs Jev avg -- Jex loses zero-shot
8 per label68.6 vs 80.2Jex avg vs Jev avg
32 per label77.6 vs 80.2Jex avg vs Jev avg -- closing in
128 per label81.6 vs 80.2Jex avg vs Jev avg -- Jex ahead
Full data85.3 vs 76.0Jex avg vs Jev avg -- Jex ahead
Instant (zero-shot, opt-in: off)53.8Jex avg, same 13 sets, label names only, no second opinion -- measured 2026-10-01
Instant + second opinion (opt-in: on)62.2Jex avg, same 13 sets -- +8.4pts over instant alone; send second_opinion:true to turn it on

Both instant rows use the same 13 datasets as the table below, measured the same way, 2026-10-01. Second opinion sends only the specific low-confidence texts it checks to an open-model provider (DeepSeek, via our inference gateway) -- off by default, opt-in only. Docs.

Dataset0832128full
AG News (topics, 4) · research licence, pending legal81.8Jev 86.881.2Jev 86.886.0Jev 86.888.8Jev 86.895.0Jev 86.8
Emotion (6) · research licence, pending legal48.0Jev 62.339.0Jev 62.350.2Jev 62.357.5Jev 62.394.0Jev 62.3
GoEmotions (28, multi-label)23.4Jev 24.9———56.4Jev 24.9
Banking77 (intents, 77)72.5Jev 88.283.8Jev 88.291.5Jev 88.293.0Jev 88.293.5Jev 88.2
CLINC150 (intents, 151)65.0Jev 71.077.5Jev 71.083.8Jev 71.087.8Jev 71.088.8Jev 71.0
TREC (question type, 6)36.8Jev 87.866.2Jev 87.883.5Jev 87.889.5Jev 87.896.5Jev 87.8
SST-2 (sentiment, 2)85.8Jev 97.085.5Jev 97.088.2Jev 97.088.2Jev 97.093.5Jev 97.0
Yahoo Answers (topics, 10) · research licence, pending legal49.0Jev 67.550.7Jev 67.562.7Jev 67.563.0Jev 67.572.5Jev 67.5
MASSIVE English (60)61.3Jev 83.071.5Jev 83.081.8Jev 83.085.8Jev 83.089.5Jev 83.0
MASSIVE German (60)58.5Jev 79.063.5Jev 79.075.0Jev 79.082.8Jev 79.081.5Jev 79.0
MASSIVE Spanish (60)57.8Jev 80.068.0Jev 80.078.8Jev 80.081.5Jev 80.082.0Jev 80.0
MASSIVE Japanese (60)61.3Jev 82.869.2Jev 82.877.5Jev 82.881.2Jev 82.883.8Jev 82.8
MASSIVE Chinese (60)55.8Jev 77.567.2Jev 77.572.8Jev 77.580.2Jev 77.582.5Jev 77.5

Honest picture: Jev leads at no examples or only a handful per label on almost every set, and leads at every size on SST-2 sentiment specifically. Jex catches up around 32-128 examples per label on most sets and leads on most sets with full training data. Instant mode (above) trails this table's "0 examples" column on average (53.8/62.2 vs 58.2) -- it's a different technique (open-weight synthetic generation vs. the original closed-model read) and not directly comparable cell-for-cell; shown honestly as its own line rather than merged into the table.

Latency: Jex p50 12.7 ms server time per classification (measured on the serving CPU, trained/ready-made head, single text); Jev's stored per-item latency on these runs was ~2-230 ms depending on the call shape, not independently re-measured here. Price: Jex $0.005-0.01 per 1,000 texts depending on mode (instant vs. ready-made/custom-trained; second opinion, opt-in, +$0.01/1,000 texts it actually checks); Jev (TypeSafe)'s published list price is $0.042 per million input tokens with output free, which works out to roughly $0.002-0.008 per 1,000 calls depending on text length -- cited from TypeSafe's own pricing, not a new call. Cells show Jex's score over Jev's (accuracy; micro-F1 for GoEmotions). Columns are examples per label; "0" means label names only. Measured 2026-09-28, stored numbers, no new Jev calls. Methodology

Ready-made models

Select by name with {"classifier": "jex/..."} -- no training, no warm-up. Each trained once on licensed public data; numbers measured 2026-10-01.

ModelTrained onHeld-out accuracy
jex/spamSMS Spam Collection (CC BY 4.0) + synthetic emails98.1% own test; 81.2% / 66.9% spam recall on Enron-Spam -- a domain-tuned eval, not held-out (the synthetic styles were chosen by looking at Enron's misses)
jex/intentCLINC150 (CC BY 3.0), full official split, 151 labels incl. out-of-scope98.3% accuracy, held-out test
jex/intent-bankingBanking77 (PolyAI, CC BY 4.0), full official split, 77 labels94.3% accuracy, held-out test
jex/sentimentSynthetic (open-weight model, 8 review domains) -- no commercially-licensed sentiment corpus found85.75% on SST-2, eval only, never trained on
jex/topicSynthetic (open-weight model, 4 news categories) -- same reasoning as sentiment82.0% on AG News, eval only, never trained on

Full numbers, model cards and licence notes: docs.

Pricing

Free

$0
  • 10,000 classify texts/month
  • 3 classifiers, 10,000 examples each
  • Report card on every model

Pay as you go

$0.005-0.01 / 1,000 texts
  • Instant classify (label names only): $0.005/1,000 texts
  • Ready-made or custom-trained classifiers: $0.01/1,000 texts
  • Training: first 10,000 examples free, then $1 per 10,000
  • Second opinion (opt-in, instant mode only): +$0.01/1,000 texts it actually checks -- sends those specific texts to an open-model provider
  • Up to 100 classifiers

Custom / Enterprise

Talk to us
  • Higher rate limits
  • Dedicated capacity
  • Fine-tuned models for hard labels

Privacy

  • Your examples train your classifier only. They never train another customer's model or the shared encoder.
  • Training texts are discarded when training finishes, unless you ask us to keep them for retraining.
  • Classification texts are not stored.
  • By default, your texts never go to a third party. The one exception is opt-in: turning on second_opinion sends the specific low-confidence texts it applies to (not your full batch) to DeepSeek, an open-weight model, through our inference gateway -- off unless you explicitly set second_opinion: true.
  • Delete a classifier and everything it holds with one call, and get a receipt.
  • Export your classifier's weights at any time.
  • A classifier unused for 90 days is deleted automatically.

Questions

How many examples do I need?

Two per label is the minimum. For a trustworthy report card, aim for 25 or more per label; accuracy keeps improving into the hundreds.

Can I start with no examples at all?

Not yet on its own. Today Jex needs labelled examples. A cold-start mode for label names only is in progress.

Which languages work?

The encoder covers more than 50 languages. You can train in one language and classify text in another, though accuracy is best when your examples include the languages you expect.

How accurate will my classifier be?

It depends on your labels and examples, which is why every model comes with a report measured on examples it never saw during training, plus the confidence level at which it is right 90% of the time.

Can one text have several labels?

Yes. Send labels: [...] instead of label and Jex trains a multi-label classifier with its own threshold.

What happens to my data?

See Privacy above. In short: used for your classifier only, discarded after training by default, deletable with a receipt.