AI running cost
Lean costs about $2 a delay; the deeper design the rulebook allows, $8 to $11; the tool-using designs are not allowed at all.
Three words first. A token is the unit AI is billed by, roughly three-quarters of a word. The lean design has ordinary code do the heavy work and asks the AI to read small packets. A deeper (multi-agent) design sends many small AI calls at once and has code check every answer. The cache lets the AI re-read text it has just seen at a fraction of the price.
A 9,000-activity schedule is about 800,000 tokens; 25,000 activities (about 2.2 million) does not fit any Claude model's working memory. So "code computes, the AI reads small packets" is not only cheaper: it is the only way large schedules work. estimate
Cost per delay event (build, round re-runs and documents)
| Design | 9,000 activities | 25,000 activities | Fits today's $5 / $3 caps? |
|---|---|---|---|
| Lean, tiered models (Haiku, Sonnet, Opus) | $1.49 ($1.17 to $1.90) | $1.93 | Yes, always |
| Lean, Opus 5.5 medium (the specification's default) | $2.23 ($1.79 to $2.79) | $3.03 | Yes, always |
| Lean, Opus 5.5 high | $3.50 ($2.72 to $4.53) | $4.70 | Yes at 9,000; 89% at 25,000 |
| Deeper, Sonnet workers under an Opus lead (no model-held tools) | $7.96 ($6.07 to $10.41) | $14.06 | Build 98%, round 27% at 9,000 |
| Deeper, Opus 5.5 high (no model-held tools) | $11.41 ($8.74 to $14.94) | $20.28 | 3% at 9,000; never at 25,000 |
| Not allowed: tool-using lead agent, Opus 5.5 high | $62 ($36 to $110) | $86 | Model-held tools |
| Not allowed: the same on Fable 5.1 | $148 | $205 | Retention |
| Not allowed: naive agent pulling the schedule in | $223 | $600 | Model-held tools |
The deeper design that obeys the rulebook
- Code builds every candidate option the rules allow and slices them into packets.
- Each packet goes to a worker in one structured call with no tools.
- Code builds and engine-tests the workers' suggestions for a second wave.
- One lead call picks, ranks and explains up to three options; critic calls return fixed objections that code checks.
- The link judge votes five times instead of three; three document writers split by audience.
No model calls another model. It costs a fifth of the tool-using design because no call re-reads a growing conversation. Whether several waves of AI-suggested changes stay within the rulebook's "the AI may suggest further changes; code builds and checks every suggestion" is a reading for the founders.
Caps it would need (founders' decision, set only in the specification): about $9 build and $10 round at 9,000 activities on Opus high (99 events in 100); about $16 and $21 at 25,000. estimate
Blind cross-check: about $35 a pass and three passes, about $105 a delay event for its tool-using design, above our $62 for the same kind; both are the kind the rulebook does not allow.
Bounding a month
Per-run caps bound each run, not the number of runs, and today's monthly counter (50 analyses a project, a placeholder) alerts and does not block. checked
- An allowance of included deeper runs per band with paid overage. Each included deeper run adds at most about $12 over a lean one, so the allowance size barely moves margin; to earn 80% on the overage alone it would sell at about $100 a run.
- A hard block at a monthly AI budget, which needs a rulebook change.
Testing cost is a budget fact
The rulebook lets no model, prompt, engine or shared-rule change go live until the graded test set passes 10 runs. checked For the tool-using design that is about $31,000 per change, against $1,000 to $1,750 lean; for the compliant deeper design it is lower but not yet costed. At pre-seed this argues for lean through the pilots. estimate
Munim's cost dashboard, checked internal: not for applications
It reproduces both cost models exactly at the central case and every price is current. In its costlier case it understates infrastructure by about a third (one activities input feeds both models); it has no multi-agent design and no routine AI; one of its options relies on a confidence source the rulebook forbids; and its headline costs carry $25 of hosting where the hosting model gives $1 to $10 at the central case. Engineering detail for the scheduler, never application text. high confidence
Measure first
- Tokens per run on a realistic schedule (a stub test rig, a few hundred dollars).
- Cache-hit rate.
- Whether a deeper design beats lean on quality at all.
- Delay events per project per month.
Margin, and the route to software margins
People time per project sets the margin. Bands C and D clear software norms; band B needs productised onboarding; band A does not work in year one.
Gross margin per project-month, by people cost to serve
One delay event a month, lean design (the specification's default), AI at the likely case, 9,000 activities; payment fees not included. Each cell: same-numbers rule / converted rule.
| People cost a project-month | A $1,500 / $1,681 | B $4,000 / $4,482 | C $8,000 / $8,965 | D $15,000 / $16,809 | Pilot $4,000 |
|---|
Reading the table
- The deeper design changes these by at most a point at one event a month, and by 3 to 8 points at 15 events a month (band B at $600: 80% / 82% instead of 84% / 85%). The AI choice moves margin per event by about $9; people cost moves it by up to $1,900.
- The pilot stage is where the company will be through YC and StartX: read the $2,500 row. 85% at band B is a claim for when onboarding is productised.
- What 75% to 80% needs: C and D get there at every people cost modelled except C at $2,500; B and the pilot need people cost at or under about $600 to $1,100 a project-month; A needs it under about $250 to $400, below anything modelled. No AI lever closes band A's gap.
What others report
- Procore, the closest public comparable: 80% gross margin (84% adjusted) in 2025. high confidence
- AI products averaged 45% in 2025, 53% projected for 2026 (ICONIQ). high confidence
- Business software median about 75% to 80%. medium confidence
- Seed investors accept about 60%+ now with a credible path to 75% to 80%; what loses them is margin that falls as usage rises. judgement
The route to software margins, largest saving first
Starting from the worst allowed case: the deeper design on Opus high for all 15 events a project-month, people $2,500, about $2,690 a project-month.
| Step | Saves a project-month |
|---|---|
| People cost $2,500 to $1,250 | $1,250 |
| People cost $1,250 to $600 | $650 |
| Send only 3 of 15 events to the deeper design | $110 |
| Medium effort and cheaper models where the test set allows | $14 |
| Sonnet workers | $10 |
| Warm caching | $2 |
"AI gets cheaper", honestly
Plan on the price per token of the tier we use staying flat: top-tier list prices have been flat to rising since 2025, a dearer tier appeared above it, and the newer way of counting tokens adds about 30%. What falls, about 3 times a year planned (Epoch AI measures 5 to 10 times), is the cost of a fixed quality bar, and only if each step is re-tested and moved to a cheaper model. Never price on the decline; treat it as margin.
Unmeasured, and not levers
- Also unmeasured and able to move small-band margin: hosting while there are only one or two customers (the whole platform floor, about $100 a month, falls on the first one); routine AI spread (chat and documents, central $6.50, about $30 at the 90th percentile); payment fees and test-set runs (not modelled).
- Not a margin lever: distilling a cheaper model from customers' approved answers, or pooling customer rules. Both need written customer permission under the rulebook. checked