· via Hacker News – Front Page (native)
Rampart is a 14.7 MB open-source browser filter that redacts PII before it reaches AI chatbots
NDS has open-sourced Rampart, a browser-native tool that strips names, SSNs and other personal data from chatbot prompts on-device using regexes plus a 14.7 MB language model.

An on-device filter for chatbot prompts
A team called NDS has released Rampart, an open-source system that redacts personally identifiable information inside the browser before a message is sent to an AI chatbot. According to the announcement posted on ndstudio.gov, a .gov address, everything runs on the user's own machine: the tool inspects text in the gap between typing a message and sending it, with no server participating at any point.
The pitch responds to a habit most chatbot users have developed without thinking. A request to tidy up an email carries your name and a colleague's; a question about a medical bill carries an address and an account number. Once submitted, that text travels to infrastructure the sender cannot audit. The team's stated design principle is that the only data you can be genuinely confident about is data that never leaves the device.
How the redaction works
Rampart combines two detectors. The first is deterministic: regular expressions backed by real validation checks catch structured identifiers such as Social Security numbers, credit cards, phone numbers, bank routing and account numbers, email addresses, IP addresses and government IDs. The second is a compact language model, MiniLM, which handles categories that regexes cannot enumerate — people's names and street addresses — by judging them in the context of the sentence.
The reported footprint is small: the whole pipeline, including the tokenizer, is 14.7 MB, and median in-browser latency is 3.9 ms on WebGPU.
The announcement's worked example shows a sentence containing a name and a Social Security number being rewritten with labeled placeholders such as [GIVEN_NAME], [SURNAME] and [SSN], while non-personal details like a salary figure pass through untouched. The browser temporarily keeps the removed values on the device so the placeholders can be filled back in afterwards.
Self-reported benchmarks
The model was trained on AI4Privacy's OpenPII 1.5M dataset plus a synthetic generator reinforcing 17 entity types with messy, conversational-style examples. It was scored end-to-end by the shipped pipeline on a 30,000-row held-out OpenPII slice spanning seven Latin-script languages, with the following private-term recall figures:
- Rampart (rules plus model, 14.7 MB): 98.42%
- GLiNER small v2.1 (model, roughly 600 MB): 94.2%
- A community BERT-small PII model (roughly 29 MB): 81.5%
- Microsoft Presidio (rules plus model, roughly 13 MB): 65%
- AWS Bedrock Guardrails (cloud model): 63.8%
These are the project's own numbers and have not been independently verified. The announcement also reports recall only, without precision or false-positive rates, which matter for a tool that rewrites outgoing text.
Why local and small matters
The announcement lays out two frustrations with existing approaches to PII removal. First, privacy claims from remote AI services are effectively unverifiable: a newly deployed runtime could begin logging sensitive input, and services carry insider threats and zero-day risk. Second, capable PII models tend to be large. The post cites OpenAI's Privacy Filter at roughly 2.8 GB, which it estimates would take about 38 minutes to download on a relatively poor 10 Mbps connection. A model under 15 MB changes where such a filter can realistically live — namely, inside a browser tab.
Current limits
Rampart is explicitly an alpha, and the team frames it as an initial safeguard rather than a complete strategy for managing personal information in AI chat. It supports seven languages, all Latin-script: English, Spanish, French, German, Italian, Portuguese and Dutch. The model is available on Hugging Face, an npm library is offered for integration, and the team has published a whitepaper describing the methodology.
Why it matters
Rampart's most interesting move is architectural rather than algorithmic. It relocates the privacy boundary from a vendor's policy to the user's own hardware, and because the code is open source, that claim can be inspected rather than taken on faith. The size and latency figures suggest on-device ML redaction has become cheap enough to be a default part of chat interfaces rather than a heavyweight download. Real caveats remain — single-source, recall-only benchmarks and alpha status among them — but as chatbots absorb more of daily computing, stripping personal data before submission looks set to become basic hygiene.
- #privacy
- #open-source
- #pii-redaction
- #on-device-ai
- #webgpu