A junior answers with a number. This answers with a range, the assumption that drives it, and the smallest experiment that would collapse it. Every rate below is an input you change in front of the client — the structure is the product, not the constants. How it calculates ↗
Cost does not scale linearly across these. It roughly triples.
This is usually the single largest cost driver, and the one clients have never been asked about.
Range comes from a 1.5×–2.0× overhead multiplier on the modelled subtotal, covering rework, failed batches, edge-case handling and coordination. It is wide on purpose: before a pilot, a narrow range is a guess wearing a suit.
Each bar is how much the total moves when that assumption is wrong by a plausible amount. The two groups are not the same kind of problem, and conflating them is how scoping conversations go wrong.
Decide — ask, don't measure
Nobody discovers these in a pilot. They are the client's call, and they belong in the first conversation rather than the third.
Measure — this is what the pilot is for
Every one of these is knowable in two weeks and unknowable from a meeting. The top bar is what the sample must pin down first.
Quoting the full corpus now prices your own uncertainty and the client pays for it. A -document sample, run end to end through the real pipeline, replaces the widest assumptions with measurements. The overhead multiplier drops from 1.5–2.0× to roughly 1.15–1.3×, and the range collapses with it.
None of these are quoted prices. They are placeholders shaped to be the right order of magnitude so the model runs. Replace each one with your own figure before it reaches a client, and argue about the assumption rather than the total.
Built as a scoping aid. It produces a defensible range and a pilot plan — not a quote. Source, method and limitations: github.com/sechan9999/doc-scoping-estimator ↗