Introducing Cipher
Survey software has been on the internet for twenty-five years, so why do we still take the data on faith?
This has been the question behind Surbee from the first day. A survey is only as good as the people answering it, and today a meaningful share of those "people" are scripts, click farms, and respondents pasting in whatever a chatbot wrote for them. The industry's answer has been a speeding rule, an attention check that says "select Strongly disagree", and an analyst cleaning the export by hand a week later. It became obvious to us that something really big was missing.
Today we are introducing Cipher: the engine that decides whether a survey response is worth keeping. Cipher reads everything a response leaves behind, from pointer movement and keystroke rhythm to the device, the network, the answers, and the way each answer was written. It fuses all of it into a calibrated fraud probability with the evidence attached, and it does that in 54 microseconds.
Think of Cipher as a lie detector that runs on every submission: raw behavior in, a probability, a confidence, and a reason out.
Extraordinary claims require extraordinary evidence, so see below for the receipts.
Defenses, Old and New
| Speeding rules, attention checks, paste bans | Cipher | |
|---|---|---|
| Built on | One hand-written threshold per rule, applied the same way to everyone. | One engine that weighs every signal against every other and updates its belief as evidence arrives. |
| Reads | One number (total time) or one answer (the trap question). | Timing, pointer, keystrokes, touch, scroll, focus, device, network, the text, and how the text was typed. |
| Outputs | Pass or fail. A fast honest person and a bot get the same verdict. | A fraud probability, a confidence interval, a risk band, and every signal that moved the score. |
| Speed | Hours to days. Someone cleans the export after fieldwork closes. | 54 microseconds for a whole response, median. Scored before the thank-you page has rendered. |
| Cost | Every extra trap question is paid for in respondent time and drop-off. | Nothing to the respondent. It runs in the background. Tiers 1 and 2 need no network at all. |
| Honest people | Rejects anyone fast, and anyone who drafted an answer in a notes app. | Flagged 0 of 5,000 honest respondents in our battery. |
| Confidence | None. A rule fires or it does not. | Every score carries an interval. Three signals and nine signals produce different confidence, and Cipher says so. |
| Use cases | Catching the laziest bots. | Bots, click farms, speed-runners, chatbot paste, AI text injected by extensions, fraud rings. |
How it works
A small tracker in the respondent's browser records how the survey is answered: 500 to 2,000 events for a typical response. When the response is submitted, Cipher reads all of it at once and returns three things: a probability that the response is fraudulent, how confident it is in that number, and the evidence that moved it.
No single signal can convict. A fast response is not fraud. A fast response with pasted answers, a failed trap question and a datacenter IP is a different matter, and telling those two apart is exactly what Cipher is built for. Strong evidence moves the score a lot, weak evidence moves it a little, and evidence in the respondent's favor pulls it back down.
Evidence
We love skeptics, and we are skeptics ourselves.
There are some claims you can easily verify:
- Speed. Every Cipher timing in this post comes from one command,
pnpm cipher:showcase, run on a laptop. - Tests. The Cipher suite runs every tier offline in a third of a second. 69 of 69 pass.
- No black box. Every flag Cipher raises is a named check with a sentence a researcher can read. You can disagree with a verdict, and you can always see why it was reached.
For our bolder claims, we want to provide as much nuance as we can.
Speed
- 0.46 microseconds to scan an answer for invisible characters and chatbot artifacts.
- 9.7 microseconds to replay how an answer was written, keystroke by keystroke, and decide whether it was composed, copied out from another window, pasted, or never typed at all.
- 54 microseconds for a full tier 3 pass over a whole response: every timing, behavioral, content, device, forensic and writing-process check, fused into one score. The 99th percentile is 98 microseconds.
- 17,386 responses per second on a single laptop core. That is about 1.5 billion responses a day, before you add a second core.
Nothing else does Cipher's job, so the fairest comparison is the closest thing the web already uses to tell humans from bots: a hosted bot check. Before one of those can say anything, your server has to send it a verification call and wait for the answer. From our office, that call alone took a median of 56 milliseconds to Cloudflare Turnstile and 450 milliseconds to Google reCAPTCHA. Cipher has the whole verdict about 1,000 times sooner than the faster of the two gets back to you, and it reads far more than a bot check ever sees.
And this is the slow version. These numbers come from one core of a laptop. Cipher's path has no network hop and no waiting on anything outside its own process, so it runs exactly as fast as the processor under it. Every server upgrade makes it faster, and every extra core adds another 17,000 responses a second.
Nuance
- Cipher timings are from an Apple M4 laptop running Node 24, on one core.
- Network lookups (IP reputation, fraud rings, cross-survey history) were switched off for this run. They run when a project's tier enables them, and they cost what a network call costs. The core verdict never has to wait for them.
- The bot-check numbers are only the server-side verification call, measured from our office with connection setup included. The challenge the respondent sees in the browser takes longer. Those services also answer a narrower question than Cipher does: whether a browser is automated, with nothing about the answers.
Head to head with the old defenses
We built a battery of 11 kinds of respondent, 1,000 of each, and scored all 11,000 with Cipher at tier 3. Five kinds are honest people with habits that trip blunt rules. Six are fraud, from crude headless bots to someone retyping a chatbot's answer by hand. Then we ran the two most common defenses over the same responses: a speeding rule that removes anyone finishing in under half the median honest time (79 seconds here), and a paste ban that removes anyone who pasted an answer.
This is the chart we care about most. The speeding rule throws out every fast but honest respondent, plus about one in twenty careful ones who happened to be quick. The paste ban throws out every person who drafted an answer somewhere else and pasted it in.
Cipher flagged 0 of 5,000 honest respondents and sent 5 of them, 0.1%, to review, where a person looks before anything happens. It saw the notes-app paste on all 1,000 of those responses, weighed it as one weak signal among many, and kept every one of them.
- Bots: 100% flagged. Headless browsers, and a stealth bot that fakes mouse movement. The stealth bot is caught on the free tier, by robotic typing and suspiciously even pacing.
- AI text injected by a browser extension: 100% caught. The answer contains invisible characters no keyboard produces, and it appeared without the keystrokes that would write it.
- Speed-runners: 100% caught. Too fast for the length, with minimal written effort.
- Chatbot paste: 96.1% caught. Pasted wholesale, and most of it carries assistant phrasing or markdown.
Overall Cipher caught 82.7% of the fraud with 100% precision: everything it flagged was fraud.
Nuance
- One row we did not catch. AI text retyped by hand reached review 0.2% of the time. The writing-process detector saw it on 992 of 1,000 responses, because someone copying from another window pauses in the middle of words to look back, where someone composing pauses between them. That evidence is attached to the response, and on its own it moves the score to 0.26. We calibrated Cipher so that no single weak signal can condemn a person. Cipher's deeper text analysis at higher tiers, switched off here, is built to carry this case, and we will publish its numbers once we have measured them.
- "Caught" means the response landed in review or above. "Flagged" means it crossed 0.6.
- The battery is generated data, and the archetypes were written by our own team, so some bias could exist. Real respondents are messier than any generator. These numbers show that each detector fires on the behavior it targets and stays quiet on honest behavior. They are not a field accuracy claim, and you should discount any vendor who gives you one without naming the data behind it, including us.
Calibrated scores
A fraud score is only useful if it means something between 0 and 1. An earlier version of our model put 97.4% of its scores exactly on 0 or 1, a step function wearing a probability's clothes. We rebuilt the scoring around evidence instead. On this battery 0% of scores are pinned. Honest respondents sit near the 0.15 base rate, ambiguous fraud lands in review, and bots sit above 0.9. A 0.43 and a 0.96 now say different things, and Cipher tells you how sure it is about each.
Nuance
- Calibration on generated data is a necessary condition, and real calibration needs labeled field data. Reviewers on live studies are building that set now, and we will publish the curve, including any part of it that looks worse than this one.
What's next
We are still in Cipher's early days, and we have a lot more in the pipeline.
- Field calibration. Every reviewer decision on a live study becomes a label, and the curve gets published.
- Deeper text analysis, measured. The retyped-AI case, with the same harness and the same charts.
- Mobile. A respondent on a phone leaves fewer signals than one on a laptop. The touch pipeline is built, and its battery is next.
- The Cipher SDK. The same engine, outside Surbee, for any form or survey tool that wants to know who is really answering.
We started Surbee because we believe research deserves data it can trust. Cipher runs on every survey you publish with us, today. Tell us what you want it to catch next.
FAQ
Does Cipher slow down my survey?
No. Cipher takes 54 microseconds and runs after submission, so your respondents never wait on them.
Will it throw out my honest respondents?
Cipher deletes nothing. It scores, flags, and explains. In our battery it flagged none of 5,000 honest respondents and sent 0.1% to review.
What happens to a flagged response?
It stays in your results with its score, its risk band, and every signal behind it. You decide whether to keep it.
Can I reproduce these numbers?
Yes. The harness is pnpm cipher:showcase. It writes the full report, the per-archetype tables, and the timings behind every chart.
What does it cost?
Tiers 1 and 2 run entirely in process. Higher tiers add network lookups and deeper text analysis, which Cipher only runs when the project's tier asks for them.