Key findings
- Jev, a fast decision model trained for calibrated probabilities, beat every general-purpose model I tried on large environments overall, at a fraction of the Claude models’ cost and latency. The exception is one messy environment, where a cheap Claude second opinion on its “maybe” pairs closes the gap.
- It also beat traditional entity resolution. Classic record linkage came close on large, regular directories (F1 0.91 vs 0.95) but collapsed on a small, messy one (0.70 vs 0.92).
- Understanding beats string matching on messy data. Jev appears to read what a similarity means (a numbered sibling account, a shared test mailbox) and links on conventions and world knowledge (nicknames, first-name logins, admin prefixes) where no identifier matches. My traditional baselines, given the same fields, couldn’t.
- Opus 5.5 fixed the same wrong merges as Jev and explained every decision, but at about 140× Jev’s cost it isn’t a viable option in large environments.
- Test it on your own data. On this task Jev’s raw probabilities looked underconfident, and independent tests on other classification tasks have put it level with mid-price LLMs and behind the best ones.
In this post
The problem
Jev is making the rounds online, particularly in the security world, where a lot of problems are classification problems. The holy grail of classification in security is “is this activity malicious or benign?”, but that problem is very noisy, so I was curious to test Jev on an easier, better-defined problem that identity security tools need to deal with: identity resolution.
For those unfamiliar with the problem, every employee in a company has a dozen accounts or more: a Google or Microsoft identity, Okta, Slack, GitHub, an AWS IAM user, a Salesforce login, often an on-premises Active Directory account and a personal Gmail address invited as a guest somewhere.
Some of these identities are federated with each other, so you can resolve them to an individual through federation. More often than not, though, federation isn’t there, and you have to deal with the real world: misspellings, account changes, abbreviations and more. So one of the jobs of an identity security or IGA product is to answer a deceptively simple question: which of these accounts belong to the same human?
jane.doe@company.examplejane.doe@company.exampleDoe, JaneJDOEjdoe-devjdoe.personal@gmail.comThis is classic entity resolution, and it has always been the kind of
problem where rules get you 80% of the way and the last 20% is a long
tail of exceptions. People write their names as “Doe, Jane” in one
system and JDOE in another. Contractors come back as B2B
guests under a new vendor’s domain. Service accounts in different AWS
accounts share a name. LLMs are a natural fit for that tail.
The reasons for picking this specific problem are somewhat mundane:
- I have a manually linked environment to check results against
- The input is bounded: one account and 20 candidates fit comfortably in a single request
- The data is private, so there is little chance of training-set contamination
- It’s also a case where speed and cost are important factors beyond just accuracy
- It is easy for a human to review the output and spot errors
This post is a test drive. Jev is TypeSafe’s model: instead of generating text, it answers yes/no questions about a piece of state with a probability, and it is meant to be calibrated. That is an appealing shape for security work, where you want a number you can threshold, queue for review or alert on, not a paragraph.
Identity resolution is a good first task to try it on. It happens at scale, mistakes are expensive in both directions, and there is a decades-old non-LLM literature to compare against. One caution up front: TypeSafe hasn’t disclosed Jev’s size, published its architecture, or reported a calibration error on any dataset, so this is a black-box test (independent review of the evidence so far).
I wanted to answer a few questions:
- Is a general-purpose LLM behind a prompt the right tool, or does a specialized model trained to output calibrated probabilities do better?
- Can a Claude model match or beat the specialized model at similar cost and speed?
- How consistent is Jev? If you ask twice, do you get the same identities?
- Do you need an LLM at all, or does traditional entity resolution do the job? And if the model wins, why?
The setup
The matching pipeline has three steps:
- Exact email. Accounts sharing the exact same email address are grouped without any AI.
- Shortlist. Every remaining account gets a shortlist of the 20 most similar-looking accounts in the environment (weighted Jaro-Winkler over names, usernames and emails).
- Decision. A model reads the account and its 20 candidates and decides which ones are the same person. A union-find then turns the accepted links into identities.
- 1Exact emailAccounts with the same address are grouped. No AI.
- 2ShortlistTop 20 look-alikes per account (weighted Jaro–Winkler over names, usernames, emails).
- 3DecisionA model scores each of the 20 candidates: same person or not? Union-find merges the accepted links into identities.
The models I compared on step 3:
- gpt-oss-20b, “the gpt-oss baseline” below: a general-purpose open-weights model prompted to list matches with a confidence and gated at 95%. It is a typical cheap, easily self-hosted way to put an LLM behind an entity-resolution pipeline. It keeps its own prompt and fixed 95% gate throughout, so comparisons with it measure the whole pipeline, not just the model.
- Jev, TypeSafe’s model, served on Cloudflare Workers AI. It answers one yes/no question per candidate with a probability. It is designed to be calibrated: 0.7 should mean “right about 70% of the time”.
- Claude Opus 5.5, Sonnet 5 and Haiku 4.5.
- Two traditional entity-resolution baselines, covered at the end.
Every Jev and Claude run simulated a from-scratch match. No model ever saw an existing grouping.
I tracked three metrics:
- Precision: when a model links two accounts, how often is it right?
- Relative recall: of the real links any tested system found on the shortlists, how many did this model find? It is relative because links that never make the shortlist are invisible to every model.
- F1: the harmonic mean of precision and recall, a single score that penalizes both over- and under-linking.
A model trained for calibration vs. a prompted one
I started with a small, manually linked environment: about 630 active user accounts across more than 20 systems, including a deliberately messy Active Directory security lab. After the exact-email step, 438 accounts went to the models, which means 8,760 candidate pairs.
Jev separated same-person from different-person pairs very well: the typical score was around 0.75 for genuine matches and around 0.02 for strangers. But it almost never says “95% sure”. With the baseline’s 95% gate, Jev rejects most real matches. At its natural cut-off of 0.5 it is excellent.
In calibration terms, that looks like underconfidence. Jev’s links are right about 95% of the time (Table 4, below), yet genuine matches typically score around 0.75 and almost none score 0.95 or above, and even the pairs sitting right at 0.5 turn out to be real about three times in four (see the stability runs below).
Others have found Jev’s raw probabilities off in both directions, depending on the data. Anthus measured a calibration error of 0.117 on a yes/no sentiment question, partly because of a tier of deliberately arbitrary labels, and an independent review of the first week’s tests found Jev overconfident on some datasets and underconfident on others. In both, recalibrating on a few hundred labels fixed most of the error: in Anthus’s test, isotonic regression fitted on 500 labels cut the 0.117 to about 0.03 (and to 0.008 with about 5,000).
At 0.5, Jev reproduced the manually linked identities perfectly. It also fixed every clear wrong merge the gpt-oss baseline made:
Table 1: The gpt-oss baseline’s clear wrong merges in the manually linked environment
| Baseline grouped together | Jev | Opus 5.5 |
|---|---|---|
| Two different employees (one’s Google, Slack and sensor accounts inside the other’s identity) | ✓ kept apart | ✓ kept apart |
Lab siblings rob.stone and rick.stone |
✓ | ✓ |
Lab users sara.stone and anna.stone |
✓ | ✓ |
A lab user and a generic test-user |
✓ | ✓ |
| An employee and a shared test account | ✓ | ✓ |
Run as a drop-in replacement with the baseline’s exact prompt and 95% rule, Opus 5.5 fixed the same mistakes and explained every decision. It cited matching security identifiers, object ids and naming conventions. It was also more conservative: where the only evidence was a shared first name, it declined, which dropped a few links a human would accept. But its cost rules it out for large datasets: it is currently 250× the gpt-oss baseline and about 140× Jev.
Table 2: Cost and latency to match this environment from scratch
| gpt-oss-20b | Jev | Opus 5.5 | |
|---|---|---|---|
| Cost | ≈ $0.09 | $0.16 | $22.05 ($11 batch) |
| Latency per account | not measured | ~0.25 s | ~3.5 s |
Scaling out: four large environments
One environment proves little, so I repeated the Jev experiment on four larger environments. They range from about 30,000 to over 100,000 accounts each. For each environment I drew 2,500 accounts at random, rebuilt their top-20 shortlists and asked Jev about every pair: 200,000 pairs in total.
Table 3: Agreement between Jev and the gpt-oss baseline’s links (Cohen’s κ)
| Org | Links both make | Only Jev | Only baseline | κ |
|---|---|---|---|---|
| Org B | 1,940 | 189 | 503 | 0.84 |
| Org C | 421 | 120 | 78 | 0.81 |
| Org A | 619 | 940 | 29 | 0.55 |
| Org D | 85 | 73 | 106 | 0.49 |
Jev ranked the baseline’s links near the top almost perfectly (AUC 0.97–0.99). So the disagreements are about where to draw the line, not noise. To see who was right, the blind judge scored 1,433 stratified pairs:
Table 4: Precision and relative recall per environment
| Org | Jev precision | Baseline precision | Jev recall | Baseline recall |
|---|---|---|---|---|
| Org A | 99% | 95% | 100% | 40% |
| Org B | 94% | 78% | 97% | 92% |
| Org C | 89% | 81% | 99% | 82% |
| Org D | 94% | 87% | 64% | 72% |
| All four | 95% | 82% | 96% | 71% |
Org B’s numbers depend on a policy decision: most links there join
plus-addressed test logins like jane.doe+qa@corp.example to
their owner. I count those as the same person. In the original
500-account sample, treating each test login as its own identity dropped
both systems’ precision to about 20%.
The error patterns were also interesting:
- Org A is a hybrid Active Directory shop. The
baseline missed most on-premises-to-cloud pairs,
e.g.
"Doe, Jane"in AD vsJDOE@corp.examplein Entra. The judge agreed with Jev on all 40 such pairs it checked. The baseline found about 4 in 10 real links. - Org C shows the classic over-linking failure:
svc-sync-devIAM users in four different AWS accounts merged into one “identity”, and two B2B guests who share a first initial and an email domain. Almost every link that only the baseline made (22 of 23 judged) was wrong. - Org D is the one place the baseline finds more real links than Jev (a second opinion fixes this; see the next section). Its directory is full of duplicate B2B guests: the same person invited once under a personal Gmail and again under a vendor address, or re-invited after a typo. Jev scores most of these between 0.3 and 0.5, just under its cut-off, and the judge usually accepts them with a hedge. Jev is still more precise there, and on F1 the two are roughly level.
Can a Claude model do better?
Opus 5.5 is too expensive to run on every account, so I tested the two cheaper Anthropic models: Haiku 4.5, and Sonnet 5 with thinking disabled and low effort. Both got exactly Jev’s request (same rules, same fields, same 10,438 shortlists) and returned a probability per candidate. I also swept each model’s cut-off from 0.5 to 0.95 and report its best, rather than forcing Jev’s 0.5 on it.
Table 5: Anthropic models vs Jev (best cut-off per model)
| Jev | Haiku 4.5 | Sonnet 5 | |
|---|---|---|---|
| Cost per 1,000 accounts | $0.62 | $11.67 (19×) | $28.86 (47×) |
| Latency p50 | 0.3 s | 2.4 s | 3.5 s |
| F1, manually linked environment | 0.845 | 0.876 | 0.868 |
| F1, large orgs | 0.948 | 0.894 | 0.911 |
On the large environments, which is where the volume is, Jev wins on
all three axes. The Claude models find nearly every real link, with
relative recall of 0.94–0.99, but they over-link. Two different people
who only share a first name, two guests from the same vendor, and
svc-reader3 vs svc-reader all get linked, even
though the prompt explicitly warns against each of these patterns. The
smaller models follow those negative rules less reliably than a model
trained specifically to answer yes/no questions like this one.
This is a better showing for Jev than most early independent tests. On social-science annotation, Ibrahim and Zaki found it behind the best LLM on 14 of 15 evaluation tasks, by a median 11.6 macro-F1 points, though the best model was picked separately for each task and Jev cost a median 44× less. A review of the first week’s independent tests puts it level with mid-price LLMs and behind the frontier, with speed and cost advantages that vary widely by setup. Identity resolution may simply suit Jev, but Sonnet 5 ran here with thinking disabled and low effort, and a stronger setting could narrow the gap on precision but increase the cost and time.
My cost and speed numbers are also more modest than TypeSafe’s launch post, which claims Jev is two orders of magnitude faster and more efficient than LLMs. Against Opus 5.5 I measured about 140× cheaper but only about 14× faster, and against Haiku 4.5 and Sonnet 5, 19–47× cheaper and 8–11× faster. However I’m using Jev on Cloudflare.
There are two exceptions to Jev’s lead. On the manually linked
environment, the Claude models correctly link staff’s
personal Gmail and Yahoo accounts to their work
identities, e.g. jdoe.personal@gmail.com ↔︎
jane.doe@company.example. On Org D, with its duplicate
guest invitations, Sonnet 5 alone beats Jev outright (F1 0.80 vs 0.65).
In both cases Jev leaves the real links in its 0.2–0.5 “maybe” zone.
That suggests a hybrid. Let Jev decide everything, and ask a Claude model only about the accounts that have a candidate in Jev’s maybe zone. That is 7% of accounts in the large environments (9% in Org D) and 13% in the manually linked environment.
Table 6: Jev plus a second opinion on its “maybe” pairs
| Cost vs Jev | F1, manually linked environment | F1, large orgs | F1, Org D | |
|---|---|---|---|---|
| Jev alone | 1× | 0.845 | 0.948 | 0.650 |
| Jev + Haiku 4.5 (link if ≥ 0.9) | 2.3–3.4× | 0.878 | 0.935 | 0.797 |
| Jev + Sonnet 5 (link if ≥ 0.7) | 4–7× | 0.890 | 0.945 (n.s.) | 0.860 |
The hybrid recovers the personal-account links and most of Org D’s duplicate guests, for roughly $1–4 more per thousand accounts. On Org D it lifts F1 from 0.65 to 0.86. The price is a little precision on the other three large environments, where the second opinion adds some look-alikes. Pooled, that is a wash with Sonnet and a small loss (−0.014) with Haiku. So the second opinion is worth turning on where a directory is known to be messy, not everywhere. I simulated it from recorded answers rather than running it live, and the cut-offs were picked on the first, smaller sample and not re-tuned here, so I treat it as a strong candidate, not a final result.
(Org D’s F1 here is lower than Table 4’s precision and recall imply, because the pool of real links now includes those the Claude models found.)
There is one more reason for caution. On other tasks, cascades like this have mainly saved money rather than added accuracy. Rao and Callison-Burch found gains of at most 1.5 points over the best single rubric judge, because the LLMs repeated nearly all of Jev’s most confident errors, and Ibrahim and Zaki found that routing Jev’s low-confidence items to an LLM matched or beat the LLM alone at a quarter to half of its cost. My hybrid beat both models on Org D by a wider margin, and its second opinion and judge are both Claude models, so correlated errors could be flattering it. Hand-checking the pairs it flipped would settle that.
Does Jev give the same answer consistently?
All of the comparisons above rest on single runs, so I ran Jev five times over the same 2,438 accounts.
Table 7: Jev run-to-run stability (5 runs)
| Manually linked environment | Large orgs | |
|---|---|---|
| Probability correlation between runs | 0.993–0.995 | 0.993–0.995 |
| Link decisions, Fleiss’ κ | 0.98 | 0.99 |
| Links made in all 5 runs (of those made in any) | 93% | 96% |
| F1 spread across runs | 0.004 | 0.006 |
It is very stable. The pairs that flip sit right at the 0.5 cut-off, and most of them (41 of 55) are real links. So the wobble mostly costs the occasional missed link, not a wrong merge. The F1 spread is several times smaller than the differences between models in Table 5, so those single-run comparisons hold up, at least on Jev’s side: I didn’t repeat the Claude runs. Majority-voting over several runs bought nothing.
Do you even need an LLM?
Entity resolution is far older than LLMs, so the obvious question is whether a traditional method would do. I tried the two I would reach for first, on exactly the same 20-candidate shortlists Jev saw:
- Fellegi–Sunter record linkage. This is the textbook statistical method from the 1960s, and it is what the popular open-source tool Splink implements. You compare fields (email, name, username, directory security ID, AWS account, department), and the model learns from the data how much each kind of agreement or disagreement is worth. It needs no labels.
- A gradient-boosted classifier. This is the standard supervised approach: the same comparisons as features, trained on labelled pairs. To mimic onboarding a new customer, I trained it on four environments and tested it on the fifth, rotating through all five.
Table 8: Traditional entity resolution vs Jev (F1, best cut-off per method)
| Jev | Fellegi–Sunter | Gradient-boosted trees | |
|---|---|---|---|
| Manually linked environment | 0.918 | 0.702 | 0.696 |
| Large orgs | 0.947 | 0.909 | 0.886 |
| Cost | $0.62 per 1,000 accounts | ≈ free, runs locally | ≈ free, runs locally |
On the large environments classic record linkage gets surprisingly close: within about 4 F1 points of Jev (95% CI 3–5 points), at no cost, and no data leaves your infrastructure. On the small, messy, manually linked environment it collapses.
The supervised model unsurprisingly has very different behavior depending on the environment:
- Inside familiar environments it is excellent. On the original sample, cross-validated within the same environments, it ranked pairs almost as well as Jev: AUC 0.94 and 0.98, against Jev’s 0.95 and 0.99. (Measured before the bug fix mentioned below, so indicative only.)
- On environments it had never seen, it fell apart in one case and slipped in the other: AUC 0.69 and 0.92.
The features were fine; the conventions don’t transfer. Every directory names things its own way, and a classifier trained on four of them has learned those four.
In fairness to the baselines, I built their features in about a day,
and a dedicated effort would probably do better. Two things cut the
other way, though. I fixed one bug after looking at their errors:
machine account names like 0030041A63F3 were being
normalized into identical “person names”. And the supervised model
learned from the same judge that grades it. Both of those flatter the
traditional methods, and they still lose.
If you already run classic record linkage, there is a practical middle ground. On large environments its most confident links are very precise, so it could be a free first pass that settles the easy pairs before any model call. I haven’t tested that combination yet.
Why does Jev win?
Jev returns a probability, not an explanation. What I can do is look at the pairs where it and both traditional methods disagree, and read the judge’s reasons for each verdict. Two patterns account for nearly all of the gap.
Similar-looking is not the same. In the 10,000 large-environment accounts, both traditional methods linked an estimated 500 pairs that Jev rejected and the judge ruled to be different identities (190 of them judged directly). Among those:
- Surface agreement is strong. 87% share an exact name, email prefix, username or directory ID.
- Most are in one system. 92% sit in the same system.
- Many aren’t people. About half are service or shared accounts.
Both traditional models score agreement field by field, so a matching
name wins. That is partly a limit of my setup rather than of the
methods: Splink, for example, can downweight
matches on common values, which might have caught some of these,
such as the two purchases@ mailboxes below. The judge’s
reasons are about what the agreement means:
msmithandmsmith1: the “1” exists because a second M. Smith joined.qa.test+branch_manager001@andqa.test+closer001@are plus-aliases of a shared test mailbox, so they are two test personas, not one person’s aliases.- Two SharePoint “All Company Owners” principals have identical names but embed different group GUIDs.
- Two
purchases@mailboxes belong to two unrelated companies.
Jev’s median probability on these was 0.10. It wasn’t on the fence.
Real links often share nothing. On the manually linked environment, Jev made 74 correct links that both traditional methods missed. Only 3% of them share any exact name, email, username or ID. Most cross systems, and a third involve a personal mailbox. The evidence is conventions and general knowledge:
- A GitHub login
jdoe-devand a work addressjane@company.example, in a company with one Jane. - “Setter, Nick” and “Setter, Nicholas” as guests from two companies in the same industry.
- A Snowflake user
JDOEfollowing the first-initial-plus-surname convention. - An
EA-jdoeadmin account and ajdoeddelegation account, both belonging to their owner.
None of these can be scored by string similarity unless someone writes a rule for each convention. A supervised model can learn them only if the training environments happen to share the convention, which explains its collapse on unseen environments.
This also explains why the gap is small on large environments. There, real links mostly go through hard identifiers such as synced security IDs and consistent email formats, which classic methods handle well. Jev’s edge there is mainly avoiding false links, not finding extra ones: about 500 look-alikes correctly refused, against an estimated 92 correct links that only Jev made.
Jev’s own misses are telling too. On the large environments they are
an estimated 300 pairs, 95% of them in the same system, with a median
probability of 0.39. They are one person’s secondary accounts: duplicate
B2B guest invitations, plus-addressed QA logins and admin variants. On
the small environment they are first-name-only logins that even the
judge accepts with a hedge. These are exactly the grey-zone pairs a
second opinion is for. Part of the plus-address case could move to a
deterministic pre-pass that links a plus-alias to its base mailbox, but
only when that mailbox belongs to one person: applied to a shared
mailbox like qa.test@, the same rule would create exactly
the false merge described above.
Conclusion
Before answering the questions I asked at the beginning, three caveats. First, the judge is an LLM. I checked it against the manually linked identities, but if it shares a blind spot with a model, that model looks better than it is. Opus judging other Claude models is the obvious case, and the reasons in “Why does Jev win?” come from the same judge. On rubric-judging tasks, Rao and Callison-Burch found that LLMs repeat nearly all of Jev’s most confident errors, which are exactly the errors an LLM judge could miss. Second, the large environments are sampled: 2,500 accounts each, except the stability runs, which used 500. Recall is only measured on the 20-account shortlist. Third, all of this uses one wording of the question, and Anthus found that rewording Jev’s questions moves both its accuracy and its calibration. Results may also shift as TypeSafe updates the model.
With that in mind:
Is a general-purpose LLM behind a prompt the right tool, or does a model trained for calibrated probabilities do better?
For bulk matching on this task, Jev beats the prompted gpt-oss baseline on large environments overall (F1 0.96 vs 0.76), though on the one directory full of duplicate guest invitations the two are level. It costs almost twice as much as the gpt-oss baseline, but it is 19–47× cheaper and roughly 10× faster than the Claude models. It also gives you a real knob: a probability for every candidate that you can threshold, queue for review or hand to a second opinion, at a price where asking about every pair is affordable. That knob still needs checking against your own labels, and prompted LLMs have one too. On question-answering benchmarks, the confidence chat models state in words is typically better calibrated than their token probabilities (Tian et al., 2023), and in one 19-model annotation study Jev’s raw probabilities were better calibrated than the stated confidence of 16 of the LLMs, though three frontier models did better (Ibrahim and Zaki).
Can a Claude model match or beat the specialized model at similar cost and speed?
Not on its own, except on one unusually messy directory. At 19–47× the cost and 8–11× the latency, Haiku 4.5 and Sonnet 5 still over-link on large directories. As a targeted second opinion on the 7–13% of accounts where Jev is unsure, they recover links it misses for a fraction of a cent per account. That is a big win on messy directories and a small precision cost on clean ones, so it belongs behind a per-environment switch, once a human check confirms the gain isn’t an artifact of a Claude judge grading a Claude second opinion.
How consistent is Jev?
Jev is consistent enough that a single run suffices. What varies between runs is a handful of pairs sitting exactly on the threshold, which is where a review queue belongs anyway.
Do you need an LLM at all, and why does it win?
On large, regular directories, classic record linkage gets within a few points, for free. On anything messier, and on any environment you haven’t trained on, it falls well behind. The model wins because it behaves as if it understands why two accounts look alike, and because it links on conventions and general knowledge that no feature set captures unless someone anticipated them.
So, as a test drive Jev passed. It was the most accurate option I tried on the large environments overall (on the small manually linked one and on the guest-heavy Org D, the Claude models and the hybrid did better), 19–47× cheaper than the Claude models, fast enough for bulk work, and stable run to run. The probability-per-candidate shape turned out to matter as much as the accuracy, because it gives a knob for review queues and second opinions.