English Deutsch 日本語
Innovius · ShinrAI 信頼 · semantic encryption

PII Benchmark Board

How well do PII detectors actually work? We measure our own releases and the field on the same texts, with the same scoring, and publish everything — including where we are behind.

🇩🇪 🇬🇧 🇺🇸 🇫🇷 🇪🇸 🇮🇹 🇵🇱 🇵🇹 🇧🇷 🇷🇺 🇺🇦 🇹🇷 🇸🇦 🇯🇵 🇰🇷 🇮🇱
15 locales · one model · 🇪🇺 EU-focus release, more coming soon · 🇮🇱 Hebrew is beta
Calm, private AI

Three releases on one board

ShinrAI 1.1

languagesDE · EN · JA
detection4 heads · 12 classes

The first open generation (2026). Still available as an archive.

ShinrAI 1.2

languages6 western languages
detection4 heads · 12 classes

Larger inventories, 1024-token window. Archive.

ShinrAI 1.3

languages15 locales
detection19 heads · 27 classes

The current release: identifier formats in the model, 51-class origin attribute, Japanese at its best level yet.

Outlook: ShinrAI 1.4 is in training — long-document redaction, the paperwork register, Hebrew toward full support.

"ShinrAI 1.3" columns show the open model behind our default serving settings — the number that counts. The Enterprise product layers additional rule-based detection on top (checksum validators, deterministic identifiers); that layer is exact by construction and is not part of these scores.

Our 15-locale test

Business & clinic letters — our 15-locale test
Realistic letters, invoices and clinic notes per language. Find every name, place, company and identifier.
Sehr geehrte Frau Keller, Ihr Termin in Leipzig ist bestätigt — Kundennummer KD-2024-88371.

F1 score (0–100, higher is better · exact spans). 200 texts per locale. The median of 95.0 includes the Hebrew beta locale. Azure AI PII, measured on the same texts: 53–76.

countrylanguagescoreF1precision / recall
🇯🇵JA
98.198.3 / 97.8
🇬🇧🇺🇸EN
97.797.2 / 98.1
🇧🇷PT-BR
97.397.7 / 96.9
🇪🇸ES
96.997.2 / 96.6
🇷🇺RU
96.496.2 / 96.7
🇫🇷FR
96.096.3 / 95.8
🇰🇷KO
95.595.9 / 95.2
🇮🇹IT
95.096.3 / 93.8
🇩🇪🇦🇹🇨🇭DE
94.495.3 / 93.5
🇵🇹PT-PT
94.295.3 / 93.1
🇵🇱PL
90.991.4 / 90.3
🇹🇷TR
90.190.7 / 89.5
🇸🇦AR
87.990.3 / 85.6
🇺🇦UA
86.789.2 / 84.3
🇮🇱HE beta
59.568.8 / 52.3

Country matters: the model knows what a street name looks like in France versus Spain, and what a company name looks like in Japan. Detection and replacements follow each country's conventions.

Public test sets — everyone in one harness

Each competitor runs at its best threshold; ShinrAI runs at its default settings. Scores are F1 (0–100, higher is better) with generous matching — a find counts if it touches the right text with the right type (see the glossary). ★ marks the best system per row. ⚠ marks systems that trained on that very test set.

Colors in the examples: person · place · identifier · organization
ai4privacy — everyday messages
Short chat messages and form snippets with personal data planted in them. Four languages.
Hi Maria Ortega, your parcel to 8 Elm Road, Leeds arrives Friday.
languageShinrAI 1.1ShinrAI 1.2ShinrAI 1.3LFM2.5 PII (Liquid AI)GLiNER2 PII (Fastino)GLiNER Large (Knowledgator)GLiNER Edge (Knowledgator)EU PII Safeguard (Tabularis)EU PII Multilang (Bards.ai)PII BERT (Gravitee)De-identifier (Stanford AIMI)DeID RoBERTa (OBI)
DE81.583.685.469.435.611.021.390.3 ★
⚠ training overlap likely
77.8
EN80.575.580.876.738.310.921.186.9 ★
⚠ training overlap likely
70.268.450.069.1
FR80.782.482.674.537.99.716.290.6 ★
⚠ training overlap likely
75.9
IT76.577.683.370.945.213.125.389.0 ★
⚠ training overlap likely
73.8
MultiNERD — encyclopedia sentences
Wikipedia-style sentences in eight languages. Find the people and organizations.
Clara Immerwahr studied chemistry in Breslau before joining the university institute.
languageShinrAI 1.1ShinrAI 1.2ShinrAI 1.3LFM2.5 PII (Liquid AI)GLiNER2 PII (Fastino)GLiNER Large (Knowledgator)GLiNER Edge (Knowledgator)EU PII Safeguard (Tabularis)EU PII Multilang (Bards.ai)PII BERT (Gravitee)De-identifier (Stanford AIMI)DeID RoBERTa (OBI)
DE72.878.192.1 ★62.345.625.737.958.591.3
⚠ training data undisclosed
EN79.383.294.960.945.817.328.558.796.2 ★
⚠ training data undisclosed
86.965.186.5
ES85.673.597.7 ★61.539.016.529.564.493.3
⚠ training data undisclosed
FR84.172.893.3 ★61.740.216.526.961.490.3
⚠ training data undisclosed
IT88.567.896.8 ★56.233.911.821.760.491.4
⚠ training data undisclosed
PL74.773.581.265.932.320.123.745.787.8 ★
⚠ training data undisclosed
PT81.073.394.762.738.415.428.861.096.1 ★
⚠ training data undisclosed
RU71.170.582.5 ★53.934.917.918.151.582.5 ★
⚠ training data undisclosed
Gretel — synthetic business documents
Machine-generated certificates, applications and records — the paperwork register.
ADOPTION CERTIFICATE — confirms the adoption of Urvashi Jaggi, identifier UID-PRWBO4TB.
languageShinrAI 1.1ShinrAI 1.2ShinrAI 1.3Azure AI PII (Microsoft)Presidio (open source)LFM2.5 PII (Liquid AI)Privacy Filter (OpenAI)OpenMed PII E5OpenMed Privacy FilterGLiNER PII (Urchade)GLiNER2 PII (Fastino)GLiNER Large (Knowledgator)GLiNER Edge (Knowledgator)Piiranha v1 (III)EU PII Safeguard (Tabularis)EU PII Multilang (Bards.ai)PII BERT (Gravitee)De-identifier (Stanford AIMI)DeID RoBERTa (OBI)
EN24.125.452.230.744.284.5 ★34.078.174.324.354.624.823.551.460.053.946.437.339.7
Nemotron-PII — synthetic US forms
Machine-generated US application forms and letters (NVIDIA test set).
I, Jason, live at 87 Avenida De La Estrella. Date of birth 1987-05-22.
languageShinrAI 1.1ShinrAI 1.2ShinrAI 1.3Azure AI PII (Microsoft)Presidio (open source)LFM2.5 PII (Liquid AI)Privacy Filter (OpenAI)OpenMed PII E5OpenMed Privacy FilterGLiNER PII (Urchade)GLiNER2 PII (Fastino)GLiNER Large (Knowledgator)GLiNER Edge (Knowledgator)Piiranha v1 (III)EU PII Safeguard (Tabularis)EU PII Multilang (Bards.ai)PII BERT (Gravitee)De-identifier (Stanford AIMI)DeID RoBERTa (OBI)
EN25.142.264.446.557.587.354.693.9 ★
⚠ trained on this test set
89.612.949.711.010.964.865.568.068.341.248.7
TAB — real court rulings (long documents)
Real European Court of Human Rights decisions, about 5,000 characters each. Anonymize the people and organizations.
…lodged by Mr Henrik Hasslund, represented by Mr Tyge Trier, a lawyer in Copenhagen
languageShinrAI 1.1ShinrAI 1.2ShinrAI 1.3Azure AI PII (Microsoft)Presidio (open source)LFM2.5 PII (Liquid AI)Privacy Filter (OpenAI)OpenMed PII E5OpenMed Privacy FilterGLiNER PII (Urchade)GLiNER2 PII (Fastino)GLiNER Large (Knowledgator)GLiNER Edge (Knowledgator)Piiranha v1 (III)EU PII Safeguard (Tabularis)EU PII Multilang (Bards.ai)PII BERT (Gravitee)De-identifier (Stanford AIMI)DeID RoBERTa (OBI)
EN24.432.878.348.2
⚠ API 5,120-char limit; long docs truncated
72.860.036.149.150.422.453.015.517.830.323.583.6 ★
⚠ training data undisclosed
65.466.972.6
MAPA — EU legal texts
Sentences from EU law (EUR-LEX) in six languages. Sparse — many sentences contain nothing to find, which makes it hard for every system.
Reference for a preliminary ruling from the Rechtbank Amsterdam concerning SF.
languageShinrAI 1.1ShinrAI 1.2ShinrAI 1.3Azure AI PII (Microsoft)Presidio (open source)LFM2.5 PII (Liquid AI)Privacy Filter (OpenAI)OpenMed PII E5OpenMed Privacy FilterGLiNER PII (Urchade)GLiNER2 PII (Fastino)GLiNER Large (Knowledgator)GLiNER Edge (Knowledgator)Piiranha v1 (III)EU PII Safeguard (Tabularis)EU PII Multilang (Bards.ai)PII BERT (Gravitee)De-identifier (Stanford AIMI)DeID RoBERTa (OBI)
DE22.643.460.5 ★55.938.627.813.434.212.210.315.86.310.119.443.148.9
EN36.822.936.223.419.331.728.653.4 ★13.14.810.23.24.229.048.234.836.342.530.9
ES53.729.539.028.23.543.632.573.0 ★13.96.210.42.94.354.849.436.0
FR22.441.548.642.925.132.56.133.011.28.017.35.27.618.131.744.4
IT29.516.913.311.27.034.010.436.6 ★10.01.72.81.21.330.130.311.9
PT30.138.248.748.825.828.813.443.518.09.418.56.912.026.444.252.4 ★

Strict-boundary scores (exact character spans) for every cell above ship in the raw run files.

Attack resistance

Chat-attack prompts
27 chat messages written the way people actually leak data to an AI — nicknames, hints, spread-out details — designed to slip past filters (from a 2026 privacy study).
"my colleague — let's call him R. — just moved to the office on Karlsplatz for project Bluebird…"

F1 score (0–100, higher is better), generous matching, identical texts for every system.

systemscore
Azure AI PII preview (Microsoft)
unreleased preview API — not their shipped product
88.8
ShinrAI 1.3 (Innovius.AI / EECC Labs)84.9
Azure AI PII (Microsoft)81.1
Presidio (open source)75.6
GLiNER2 PII (Fastino)74.3
GLiNER PII (Urchade)70.9
OpenMed Privacy Filter67.1
OpenMed PII E566.2
Privacy Filter (OpenAI)55.0
Piiranha v1 (III)27.2

Microsoft's shipped product (Azure AI PII GA) scores 81.1 here — below ShinrAI 1.3. Their preview API leads by 3.9; it is not generally available.

Japanese head-to-head 🇯🇵

Our name is Japanese — so is the test
300 Japanese business and clinic letters. ShinrAI (信頼) versus Liquid AI's dedicated Japanese PII model.
担当者 入迫 かつなが 様 — 川崎フロンターレ川崎市中原区

ShinrAI 1.3

98.1

F1, generous matching · grades people, places, streets and companies here; detects 27 classes overall.

LFM2 PII Extract JP (Liquid AI)

56.6

Their Japanese-only extraction model. It outputs five categories: names, addresses, companies, e-mails, phone numbers.

Long documents — how ShinrAI reads them

How long text is processed

The model reads 1024 tokens at once and slides over longer text with overlapping windows, merging spans at the seams — a 30-page file processes as one document. The serving layer adds automatic segmentation for conversational input: short text decodes whole; longer text also gets a sentence-level pass, and the two agree before a span counts.

Measured on long documents

At 4× letter length the scores hold: EN 96.6, DE 93.8 — no quality cliff. The court-ruling test above (TAB, ~5,000 characters per document) runs through this exact path. Fairness note: competitor models limited to 512 tokens were given the same windowing on long documents, so the table measures their models, not their truncation defaults.

Areas of improvement

We publish these because you should choose tools on measured numbers — ours included. All items below are on the 1.4 programme; the first is in training now.

areameasured todaystatus
Very long records (REDACT test)36.5 vs Azure 51.2A dedicated long-record training corpus is part of the 1.4 campaign.
Synthetic US-paperwork register (Gretel, Nemotron)52.2 / 64.4 partial vs Liquid AI 84.5 / 87.3A paperwork-register training track is scoped for 1.4.
Court-document span conventions (TAB)detection 78.3 partial, but strict spans 34.4We find the entities; the span boundaries follow different conventions (honorifics, full institution names). A serving-side convention profile is in work — no retraining needed.
Chat-attack prompts84.9 vs Azure preview 88.8Already ahead of Presidio and GLiNER2; decoder work for conversational text continues in 1.4.
Encyclopedic register in 8 of 15 languagesimproved for 7 languages in 1.3; the other 8 regressed on encyclopedia textBusiness-document quality is unaffected. Rebalancing both is the headline 1.4 training goal — running now.

How fast is ShinrAI on your hardware?

Time to detect PII in one request, measured with the open model innovius/shinrai-pii-m-v1.3 (full precision, ONNX Runtime). Four text sizes: a chat message (62 tokens ≈ 45 words), a paragraph (317 tokens ≈ 230 words — the size behind every ShinrAI latency figure), a page (1,024 tokens, the model's window) and a long document (10,000 tokens ≈ 15 pages). Plain numbers are measured on that machine; is an estimate from the model's compute cost (±40 %); ~ is scaled from a measured sibling machine.

MachineRAM / VRAMChat messageParagraphPageLong documentWhat it is good for
Single-board computers
Raspberry Pi 4 Model B (4 GB)4 GB · too small≈706 ms≈3.7 s≈12.5 s≈2.6 minDocuments & data pipelines only — too slow for chats
Raspberry Pi 4 Model B (8 GB)8 GB · fits≈706 ms≈3.7 s≈12.5 s≈2.6 minDocuments & data pipelines only — too slow for chats
Raspberry Pi 5 (8 GB)8 GB · fits≈193 ms≈982 ms≈3.3 s≈40.4 sAgentic AI & workflows (not real-time chat)
Raspberry Pi 5 (16 GB)16 GB · fits≈193 ms≈982 ms≈3.3 s≈40.4 sAgentic AI & workflows (not real-time chat)
Mini PCs
Intel N100 mini PC (16 GB)16 GB · fits≈174 ms≈835 ms≈2.8 s≈33.9 sAgentic AI & workflows (not real-time chat)
Laptops
Apple M1 (MacBook Air class, 4P+4E)8 GB · fits~87 ms~458 ms~1.8 s≈16.9 sAgentic AI & workflows (not real-time chat)
Apple M4 (MacBook Air / iMac / Mac mini, 4P+6E)16 GB · fits~16 ms~78 ms~310 ms~4.0 sReal-time: AI chats & assistants
x86 laptop, 8 cores (Ryzen 7 7840U / Core Ultra 7 class)16 GB · fits~48 ms~189 ms~676 ms~6.8 sAgentic AI & workflows (not real-time chat)
Desktops & workstations
Apple M4 Pro (Mac mini 2024, 10P+4E)64 GB · fits13 ms56 ms233 ms3.1 sReal-time: AI chats & assistants
AMD Threadripper 7960X workstation (24 cores, CPU only)64 GB · fits≈21 ms≈64 ms≈193 ms≈2.3 sReal-time: AI chats & assistants
Edge GPUs (Jetson)
NVIDIA Jetson Orin Nano Super (8 GB)8 GB · fits≈24 ms≈106 ms≈344 ms≈4.2 sAgentic AI & workflows (not real-time chat)
NVIDIA Jetson Orin NX (16 GB)16 GB · fits≈26 ms≈116 ms≈380 ms≈4.6 sAgentic AI & workflows (not real-time chat)
NVIDIA Jetson AGX Orin (64 GB)64 GB · fits≈13 ms≈45 ms≈140 ms≈1.7 sReal-time: AI chats & assistants
Desktop / workstation GPUs
NVIDIA DGX Spark (GB10)128 GB · fits≈6 ms≈9 ms≈19 ms≈175 msReal-time: AI chats & assistants
NVIDIA RTX 3090 (24 GB)24 GB · fits≈6 ms≈9 ms≈17 ms≈151 msReal-time: AI chats & assistants
NVIDIA RTX 4090 (24 GB)24 GB · fits≈6 ms≈8 ms≈16 ms≈131 msReal-time: AI chats & assistants
Servers & cloud VMs — CPU only
AMD EPYC 7002 'Rome' VM, 2 vCPU (cgroup limit)6 GB · fits131 ms589 ms2.2 s29.8 sAgentic AI & workflows (not real-time chat)
AMD EPYC 7002 'Rome' VM, 4 vCPU6 GB · fits83 ms350 ms1.2 s15.5 sAgentic AI & workflows (not real-time chat)
AMD EPYC 7002 'Rome' VM, 8 vCPU6 GB · fits59 ms237 ms845 ms8.5 sAgentic AI & workflows (not real-time chat)
AMD EPYC 7002 'Rome' VM, 16 vCPU6 GB · fits76 ms215 ms668 ms7.2 sAgentic AI & workflows (not real-time chat)
Cloud VM, 4 vCPU x86 (AMD EPYC 9V74 "Genoa", shared, GitHub-hosted runner)16 GB · fits105 ms514 ms1.9 s≈21.0 sAgentic AI & workflows (not real-time chat)
Cloud VM, 4 vCPU Arm (Neoverse N2 / Cobalt 100, GitHub-hosted runner)16 GB · fits92 ms442 ms1.7 s≈18.0 sAgentic AI & workflows (not real-time chat)
Cloud VM, 4 vCPU Intel Ice Lake / Sapphire Rapids (AVX-512 VNNI)16 GB · fits~70 ms~298 ms~1.0 s~13.1 sAgentic AI & workflows (not real-time chat)
Hugging Face Space, cpu-upgrade (8 vCPU, 32 GB)32 GB · fits~83 ms~350 ms~1.2 s~15.5 sAgentic AI & workflows (not real-time chat)
Servers — GPU
NVIDIA Tesla P40 (24 GB, Pascal 2016)24 GB · fits9 ms25 ms93 ms1.6 sReal-time: AI chats & assistants
NVIDIA T4 (16 GB)16 GB · fits≈10 ms≈31 ms≈93 ms≈1.1 sReal-time: AI chats & assistants
NVIDIA L4 (24 GB)24 GB · fits≈6 ms≈10 ms≈19 ms≈178 msReal-time: AI chats & assistants
NVIDIA A10 (24 GB)24 GB · fits≈6 ms≈9 ms≈19 ms≈171 msReal-time: AI chats & assistants
NVIDIA L40S (48 GB)48 GB · fits≈6 ms≈7 ms≈10 ms≈62 msReal-time: AI chats & assistants
NVIDIA A100 (80 GB)80 GB · fits≈6 ms≈7 ms≈11 ms≈72 msReal-time: AI chats & assistants
NVIDIA H100 SXM (80 GB)80 GB · fits≈5 ms≈6 ms≈7 ms≈26 msReal-time: AI chats & assistants

Raspberry Pi: works for document and data pipelines (a paragraph takes about a second on a Pi 5, four on a Pi 4), not for chats. 4 GB RAM minimum with the compact model file layout, 8 GB comfortable, 64-bit OS.

Memory: one full-precision model session needs about 2.0 GB of RAM as shipped, or about 0.9 GB with the model file in ONNX external-data layout (same speed) — the layout we recommend for small devices.

Through the ShinrAI service the default long-input mode re-reads texts over 1,200 characters sentence by sentence for better recall; that costs about 3× on a paragraph. The rows below show both modes on the same GPU.

ShinrAI service on a Tesla P40 (2016)Chat messageParagraphLetter (2 pages)Long document
model only (whole text in one pass)9 ms25 ms175 ms1.6 s
default: long-input mode (sentence pieces)9 ms69 ms348 ms4.9 s

Generated 2026-09-02 from research/hardware-map/hardware-map.json (scripts/hardware/render-map.py).

What gets detected — 19 heads, 27 classes

classShinrAI 1.3Azure PIIGLiNER2PresidioOpenAI PF
PERSON
common / uncommon / rare (name rank)
native headPersonzero-shotPERSON (spaCy)private persons only
CITY
major / medium / small (population)
native headCity/Location (preview)zero-shotLOCATION (spaCy)
STREET
generic / specific / local
native headAddresszero-shotprivate address
ORG
international / national / regional (sitelinks)
native headOrganizationzero-shot (P 15–33%)ORGANIZATION (off by default)
EMAIL
std
native headEmailzero-shotEMAIL_ADDRESSprivate email
USERNAME
std
native headzero-shot
URL
std
native headURLzero-shotURLprivate url
PHONE
std
native headPhoneNumberzero-shotPHONE_NUMBERprivate phone
ACCOUNT
std
native headIBAN / BankAccountzero-shotIBAN_CODE, US_BANK_NUMBER
WALLET
std
native headzero-shotCRYPTO
CARD
std
native headCreditCardNumberzero-shotCREDIT_CARD
NATIONAL_ID
std
native head~70 country ID categorieszero-shot~76 country recognizers
PLATE
std
native headLicensePlate (preview)zero-shot
DOB
std
native headDateOfBirth (preview)zero-shotDATE_TIME (any date)
NETADDR
std
native headIPAddresszero-shotIP_ADDRESS
POSTAL_CODE
std
native headZipCode (preview)zero-shot
CUSTOMER_ID
std
native headzero-shot
SECRET
std
native headPassword (preview)zero-shot
CODENAME
std
native headzero-shot (43% recall on the 27 prompts; 7 languages)
One record, everything at once
Dr. Anna Keller (Leipzig, Lindenweg 12) of Volksbank Nordparkanna.k@example.de · +49 170 1234567 · IBAN DE89 3704 0044 0532 0130 00 · customer KD-448291 · project Bluebird
Plus three attribute signals no other system ships — cultural origin (51 classes), gender expression, name part — the steering data for replacements that read naturally and restore losslessly.

Glossary

F1 scoreOne number from 0 to 100 combining "found everything" (recall) and "no false alarms" (precision). Higher is better.
StrictA find only counts with the exact character boundaries and the right type.
Generous / partialA find counts if it touches the right text with the right type. Fairer across systems with different span conventions.
Test set × languageEach table row is one public test set in one language, typically 300 texts.
Best thresholdCompetitors report their best score across sensitivity settings. ShinrAI always reports its default settings.
⚠ trained on this test setThe system saw this test's data during training — its score there is inflated and not comparable.
Precision / recallPrecision: of everything flagged, how much was right. Recall: of everything present, how much was found.
Beta localeHebrew ships for integration and feedback; its quality is below release level and is shown, and counted, everywhere.

Method

ShinrAI (信頼, Japanese for trust) · Innovius × EECC Research Labs · trained at the Jülich Supercomputing Centre (JURECA, WestAI) · open model weights · try ShinrAI live — PII playground