How It Works

From Your Data to a Certified Twin

The whole journey · what we do with your data, what you get back, and exactly how good it is. Written once, the same way for any process in any domain: a fermentation batch, a battery cycle, a wafer lot, a grow cycle, a treatment run.

The Journey

Your data does not have to be clean, unified, or in one place. Point us at your systems as they are · historian telemetry, a separate sensor and hardware config, offline lab results, batch logs, different sample rates, unfamiliar tag names. We handle the mess.

A pilot runs about 5 to 15 business days from your data handoff · long enough to ground and truly validate, not an instant black box. Your data is kept only for that window, then destroyed with proof.

1

Point us at your data

day 0 · as-is

Send a scoped slice as it is · historian, sensor config, lab results, batch logs, any sample rate. A missing piece never bars the work; it just flags what it changes. You do not pre-clean anything.

2

We unify & quality-check it

~1 to 2 business days

Your scattered sources become one clean, aligned, provenance-tracked dataset in your own terms · tags mapped, timestamps aligned, config attached, every gap flagged. We measure your sensor noise and flag faulty hardware.

3

We draw the honest line

same window

Before building anything, we tell you exactly what your data can support today · and the one thing that unlocks more. We never fabricate a result from data that cannot carry it.

4

We ground & certify your twin

~5 to 15 business days

We fit a fast predictor of your process and derive a proven error bound, validated on held-out runs · then deliver the frozen certified twin, its proof packet, and an offline verifier. This is the deliverable.

5

You close out · we prove deletion

on your acknowledgement

When your engagement wraps, you close out in My Engagements and acknowledge receipt. That acknowledgement is what authorizes deletion · we never delete before you accept · and you immediately get a signed Deletion Authorization receipt. We then destroy your Source Data within 5 business days and post a signed Deletion Completed receipt · together they form a closed, verifiable authorize → delete → confirm chain. Both are content-free: they record that it happened, never your data.

Is Your Data Enough?

The honest line

what your data supports today · and what unlocks more
per domain
Not enough yet
No process variable measured over time, or only a single run. We tell you the one thing to add · we do not guess.
Descriptive
Enough to unify your data and give you a data-quality and operating-regime read · not yet a predictor.
Bounded, un-grounded
Enough to build a fast predictor with a worst-case error bound and honest refusal, but no ground-truth outcome yet · bounded against our model, not yet your reality.
Grounded
Enough runs to ground against your reality. From here we state the actual guarantee as a number (next section) · and it improves as you add runs.

How many runs is enough?

computed from your data · a rule, not a guess
up front

A "run" is one complete cycle of your process · a batch, a charge/discharge cycle, a wafer lot, a grow cycle, a treatment day, a test coupon. The number of runs needed is derived from two things: enough distinct conditions to span your operating range, and enough held-out evidence to confirm the bound holds. Noisy sensors raise it.

It differs by domain · a batch is information-rich in a handful; a battery needs many cycles; a seasonal process needs seasons. The rule is the same everywhere; only the per-run richness changes. Your exact figure is computed at intake · never on our word.

How we handle your sensor noise

measured from your data · not assumed
at intake

Every real sensor is noisy, so we account for it in three places · and the first one happens the moment your data arrives:

1 · We measure it
At intake we estimate each sensor's noise floor directly from your series (using the fine-scale jitter that a real signal cannot explain), and we flag any sensor that is noise-dominated · where the noise is as big as the signal it is meant to read.
2 · It sets the bar
Noisier sensors need more runs to reach the same coverage · the required run count rises with measured noise, so we never quietly promise a clean-data result on noisy data.
3 · It lives in the bound
Normal sensor noise is inside the certified bound by construction · it is part of what the bound is fit to cover. That is exactly why routine noise never trips your twin, and only abnormal change (a failing probe, a real drift) does.

If a sensor is too noisy to be useful, we tell you up front · calibrate or replace it, or add replicate runs to average the noise down · rather than burying it in a number.

Your Two Numbers · Coverage vs. Confidence
Read it as one sentence:
"With 90% confidence, the twin's error stays inside its stated bound for at least 99% of your real operating conditions."
The green number is coverage · how often the twin keeps its promise. The blue number is confidence · how sure we are of that first number. They are not the same thing, and neither is "accuracy."

Coverage

how often the twin keeps its promise
the big number
99%

Out of all the real situations you will actually run, coverage is the fraction where the twin's true error stays inside the bound we stated. At 99%, in 99 of every 100 real conditions the error is within promise · and on the rest the twin refuses or flags rather than guessing. It never silently exceeds its bound.

Rises when
you have more runs across more of your range.
It is NOT
"the twin is 99% accurate." It is how often the error is inside the bound · the bound's size is a separate number.

Confidence

how sure we are of the coverage number
the trust number
90%

We only got to see a finite number of your runs, not infinite. Confidence is how certain we are that the true coverage is at least what we claimed. At 90%, there is only about a 1-in-10 chance we got an unlucky sample and the real coverage is lower. More data lets us raise it.

Rises when
you give us more runs · a bigger sample is harder to fool.
It is NOT
the coverage. It is our certainty about the coverage.

The everyday version

why both numbers matter
analogy

Two products, both rated 4.8 stars. One has 6 reviews, the other 6,000. The rating is like coverage · how good it is. The number of reviews is like confidence · how much you should trust that rating. A 4.8 from 6 reviews might be luck; a 4.8 from 6,000 is real.

Same with your twin: a high coverage claimed from a handful of runs would be an overclaim, so we always report the confidence alongside it. We never quote coverage without telling you how sure we are of it.

How Your Numbers Grow · And the Third One

More runs, higher coverage

computed from your run count · before we build anything
up front

The shape below is universal, at a fixed 90% confidence · your exact figures come from your data at intake.

Independent runsCoverage we can confirm
~6 runs~68%
~12 runs~91%
~25 runs~96%
~60 runs~99%
~25 runs, noisy sensors~86%

The third number: bound size

what coverage and confidence do NOT tell you
at grounding

Coverage and confidence say how often the error is inside the bound and how sure we are. They do not say how big the bound is · the actual error in your variable's own units (e.g. "within ±X on your key output"). That depends on your process dynamics and is fixed once we ground on your data.

Coverage
how often the error is inside the bound (up front).
Confidence
how sure we are of that coverage (up front).
Bound size
how large the bound is, in your units (at grounding).

A bound is only useful if it is both well-covered and tight enough to matter · we report all three so you can judge, not just trust.

What You Receive · A Tiered Ladder

Each result labeled by how defensible it is

you choose the tier · higher tiers need more data
you choose
TierWhat you receiveHow defensible
Unified datasetYour fragmented data, reconciled + quality-checked + gap-mapped, in your own terms.fact
Data & regime readCoverage, sensor health, drift already present, and where you run in and out of your own compliance ranges.fact
Certified twinA fast bounded predictor + proof packet + honest refusal · verified offline with a public key, no reliance on our word.bounded
DiagnosticWhich inputs move which outputs and by how much, what drives out-of-spec, and your headroom to your limits.model-relative
Analysis & recommendations roadmapTwin-grounded operating suggestions to hold your ranges or improve a goal, ranked by confidence · in-envelope only, an input to your judgment. A future add-on · offered once your twin is grounded and only where the data earns it.advisory

Defensibility ladder: fact (measured) › bounded (certified twin) › model-relative (from the twin) › advisory (a recommendation, your call). We never label a suggestion as a certified fact. We evaluate against your declared targets and compliance ranges · your SOP or spec · not our idea of optimal, and you keep every compliance, quality, and release decision.

Bottom line: we turn your scattered data into a clean dataset, tell you honestly what it can support, and · where it earns it · hand you a certified twin whose guarantee you can read plainly and verify yourself. Coverage is how often the promise holds, confidence is how sure we are of that, and bound size is how tight it is. We never quote one without the others · that is what "proof, not promises" means here.
Questions
What exactly is a "certified twin"?

A fast software model of your process that predicts what happens next · and, unlike a normal model, it comes with a proven worst-case error bound and refuses to answer when your process moves outside what it was checked for. It is the prediction plus the proof.

What do I actually get delivered?

A frozen certified twin you can run anywhere, a sealed proof packet (the bound and the tests behind it), and an offline verifier so anyone can check the proof with a public key · no call to us, no trust in our word.

How is this different from a normal digital twin or ML model?

Most models give you a number with no idea when to trust it. Ours gives you a number with a bound · how far off it can be, how often that holds, and honest refusal outside its range. It is auditable, not a black box.

Does my raw data leave my environment?

You choose. It can stay fully in-environment (the twin runs behind your firewall, data never leaves at runtime), or you send a scoped, non-retained slice we destroy with a signed receipt. We never pool, sell, or train on your data. See the Data Policy.

How much of my data do I need?

Usually less than you think · most established plants already have it in their historian. The exact number is computed from your data at intake, per domain (a batch is information-rich in a handful of runs; a battery needs many cycles). We tell you up front, never on our word.

What happens if my process drifts?

The twin watches its own error. Normal variation is already inside the bound; when reality moves beyond it, the twin flags it and abstains rather than guessing · that is your signal to re-certify. It never silently rots.

Which industries does this work for?

Any process with data over time · fermentation, pharma, battery, ag-bio and plant biotech, water, materials, semiconductor, controlled-environment ag, magnets, robotics. One method, your process.

How do I get started?

Try the Cleanroom with your own data (or our examples) to see an honest readiness read in seconds · then complete a pilot order on the Agreement & Order page.