Your data does not have to be clean, unified, or in one place. Point us at your systems as they are · historian telemetry, a separate sensor and hardware config, offline lab results, batch logs, different sample rates, unfamiliar tag names. We handle the mess.
A pilot runs about 5 to 15 business days from your data handoff · long enough to ground and truly validate, not an instant black box. Your data is kept only for that window, then destroyed with proof.
Point us at your data
day 0 · as-isSend a scoped slice as it is · historian, sensor config, lab results, batch logs, any sample rate. A missing piece never bars the work; it just flags what it changes. You do not pre-clean anything.
We unify & quality-check it
~1 to 2 business daysYour scattered sources become one clean, aligned, provenance-tracked dataset in your own terms · tags mapped, timestamps aligned, config attached, every gap flagged. We measure your sensor noise and flag faulty hardware.
We draw the honest line
same windowBefore building anything, we tell you exactly what your data can support today · and the one thing that unlocks more. We never fabricate a result from data that cannot carry it.
We ground & certify your twin
~5 to 15 business daysWe fit a fast predictor of your process and derive a proven error bound, validated on held-out runs · then deliver the frozen certified twin, its proof packet, and an offline verifier. This is the deliverable.
You close out · we prove deletion
on your acknowledgementWhen your engagement wraps, you close out in My Engagements and acknowledge receipt. That acknowledgement is what authorizes deletion · we never delete before you accept · and you immediately get a signed Deletion Authorization receipt. We then destroy your Source Data within 5 business days and post a signed Deletion Completed receipt · together they form a closed, verifiable authorize → delete → confirm chain. Both are content-free: they record that it happened, never your data.
The honest line
- Not enough yet
- No process variable measured over time, or only a single run. We tell you the one thing to add · we do not guess.
- Descriptive
- Enough to unify your data and give you a data-quality and operating-regime read · not yet a predictor.
- Bounded, un-grounded
- Enough to build a fast predictor with a worst-case error bound and honest refusal, but no ground-truth outcome yet · bounded against our model, not yet your reality.
- Grounded
- Enough runs to ground against your reality. From here we state the actual guarantee as a number (next section) · and it improves as you add runs.
How many runs is enough?
A "run" is one complete cycle of your process · a batch, a charge/discharge cycle, a wafer lot, a grow cycle, a treatment day, a test coupon. The number of runs needed is derived from two things: enough distinct conditions to span your operating range, and enough held-out evidence to confirm the bound holds. Noisy sensors raise it.
It differs by domain · a batch is information-rich in a handful; a battery needs many cycles; a seasonal process needs seasons. The rule is the same everywhere; only the per-run richness changes. Your exact figure is computed at intake · never on our word.
How we handle your sensor noise
Every real sensor is noisy, so we account for it in three places · and the first one happens the moment your data arrives:
- 1 · We measure it
- At intake we estimate each sensor's noise floor directly from your series (using the fine-scale jitter that a real signal cannot explain), and we flag any sensor that is noise-dominated · where the noise is as big as the signal it is meant to read.
- 2 · It sets the bar
- Noisier sensors need more runs to reach the same coverage · the required run count rises with measured noise, so we never quietly promise a clean-data result on noisy data.
- 3 · It lives in the bound
- Normal sensor noise is inside the certified bound by construction · it is part of what the bound is fit to cover. That is exactly why routine noise never trips your twin, and only abnormal change (a failing probe, a real drift) does.
If a sensor is too noisy to be useful, we tell you up front · calibrate or replace it, or add replicate runs to average the noise down · rather than burying it in a number.
"With 90% confidence, the twin's error stays inside its stated bound for at least 99% of your real operating conditions."
Coverage
Out of all the real situations you will actually run, coverage is the fraction where the twin's true error stays inside the bound we stated. At 99%, in 99 of every 100 real conditions the error is within promise · and on the rest the twin refuses or flags rather than guessing. It never silently exceeds its bound.
- Rises when
- you have more runs across more of your range.
- It is NOT
- "the twin is 99% accurate." It is how often the error is inside the bound · the bound's size is a separate number.
Confidence
We only got to see a finite number of your runs, not infinite. Confidence is how certain we are that the true coverage is at least what we claimed. At 90%, there is only about a 1-in-10 chance we got an unlucky sample and the real coverage is lower. More data lets us raise it.
- Rises when
- you give us more runs · a bigger sample is harder to fool.
- It is NOT
- the coverage. It is our certainty about the coverage.
The everyday version
Two products, both rated 4.8 stars. One has 6 reviews, the other 6,000. The rating is like coverage · how good it is. The number of reviews is like confidence · how much you should trust that rating. A 4.8 from 6 reviews might be luck; a 4.8 from 6,000 is real.
Same with your twin: a high coverage claimed from a handful of runs would be an overclaim, so we always report the confidence alongside it. We never quote coverage without telling you how sure we are of it.
More runs, higher coverage
The shape below is universal, at a fixed 90% confidence · your exact figures come from your data at intake.
| Independent runs | Coverage we can confirm |
|---|---|
| ~6 runs | ~68% |
| ~12 runs | ~91% |
| ~25 runs | ~96% |
| ~60 runs | ~99% |
| ~25 runs, noisy sensors | ~86% |
The third number: bound size
Coverage and confidence say how often the error is inside the bound and how sure we are. They do not say how big the bound is · the actual error in your variable's own units (e.g. "within ±X on your key output"). That depends on your process dynamics and is fixed once we ground on your data.
- Coverage
- how often the error is inside the bound (up front).
- Confidence
- how sure we are of that coverage (up front).
- Bound size
- how large the bound is, in your units (at grounding).
A bound is only useful if it is both well-covered and tight enough to matter · we report all three so you can judge, not just trust.
Each result labeled by how defensible it is
| Tier | What you receive | How defensible |
|---|---|---|
| Unified dataset | Your fragmented data, reconciled + quality-checked + gap-mapped, in your own terms. | fact |
| Data & regime read | Coverage, sensor health, drift already present, and where you run in and out of your own compliance ranges. | fact |
| Certified twin | A fast bounded predictor + proof packet + honest refusal · verified offline with a public key, no reliance on our word. | bounded |
| Diagnostic | Which inputs move which outputs and by how much, what drives out-of-spec, and your headroom to your limits. | model-relative |
| Analysis & recommendations roadmap | Twin-grounded operating suggestions to hold your ranges or improve a goal, ranked by confidence · in-envelope only, an input to your judgment. A future add-on · offered once your twin is grounded and only where the data earns it. | advisory |
Defensibility ladder: fact (measured) › bounded (certified twin) › model-relative (from the twin) › advisory (a recommendation, your call). We never label a suggestion as a certified fact. We evaluate against your declared targets and compliance ranges · your SOP or spec · not our idea of optimal, and you keep every compliance, quality, and release decision.
What exactly is a "certified twin"?
A fast software model of your process that predicts what happens next · and, unlike a normal model, it comes with a proven worst-case error bound and refuses to answer when your process moves outside what it was checked for. It is the prediction plus the proof.
What do I actually get delivered?
A frozen certified twin you can run anywhere, a sealed proof packet (the bound and the tests behind it), and an offline verifier so anyone can check the proof with a public key · no call to us, no trust in our word.
How is this different from a normal digital twin or ML model?
Most models give you a number with no idea when to trust it. Ours gives you a number with a bound · how far off it can be, how often that holds, and honest refusal outside its range. It is auditable, not a black box.
Does my raw data leave my environment?
You choose. It can stay fully in-environment (the twin runs behind your firewall, data never leaves at runtime), or you send a scoped, non-retained slice we destroy with a signed receipt. We never pool, sell, or train on your data. See the Data Policy.
How much of my data do I need?
Usually less than you think · most established plants already have it in their historian. The exact number is computed from your data at intake, per domain (a batch is information-rich in a handful of runs; a battery needs many cycles). We tell you up front, never on our word.
What happens if my process drifts?
The twin watches its own error. Normal variation is already inside the bound; when reality moves beyond it, the twin flags it and abstains rather than guessing · that is your signal to re-certify. It never silently rots.
Which industries does this work for?
Any process with data over time · fermentation, pharma, battery, ag-bio and plant biotech, water, materials, semiconductor, controlled-environment ag, magnets, robotics. One method, your process.
How do I get started?
Try the Cleanroom with your own data (or our examples) to see an honest readiness read in seconds · then complete a pilot order on the Agreement & Order page.