BaseMod.ai Regulatory Rehearsal
Rehearse the approval before you file.
BaseMod agents play both sides of a land review. A developer team submits the plan. Simulated road, health, zoning, soil-erosion and building reviewers object, cite the rule and recheck every fix, round by round, on the real lot.
Product
Run a rehearsal
Pick a venue and step through a review the way the engine runs it. Objections pin themselves to Lot 15 as the simulated agency raises them. You answer each one with one of four responses, the referee decides what closes, and the run ends with a result card. This is the Review Room staff would use in Land Explorer, on a daylight satellite plan of the real lot. Every scenario here is an example, not a real run.
How it works
Two teams, one engine, one referee
The pro team on the left proposes and answers. The regulator team on the right challenges, grouped the way Michigan authorities actually work: the township's own review cluster, the bodies that vote, the county and state agencies you apply to separately, and the building department. In the middle sits the machinery that keeps the contest honest. The engine owns every number and every state change. The referee decides which objections and answers count. You sit at the bottom, in the Review Room.
Select any node to open its role card beside it; Esc closes it. Every card is also in the roster tables further down. On a phone, scroll the diagram sideways.
One round
Twelve steps from a saved project to a result card
This is one review round on one venue, such as the county road commission's driveway permit. The same machinery runs for zoning, health, soil erosion and the building department. Only the rule pack and the authority change. Use Next, or press Play.
Thousands of runs
Where the thousands actually go
Your instinct to test a design thousands of times is right. The research says where to spend them. Repeated runs of the same language model are correlated, drift toward the same answers and cannot be replayed by a seed. So the thousands run in code, on seeded engines, and agents work where judgment is needed. Six layers do the work.
Rule packs: the real counter-agents
Edition-pinned, table-valued checks per authority: setbacks, sight distance, isolation distances, frontage, triggers. Each row cites a section number, a paraphrase, a source date and a hash, never code text.
Once per project revision
≈ $0Seeded Monte Carlo and discrete-event runs
Review rounds, agency clocks, inspection readiness and crews. Every run carries a seed and replays exactly.
Thousands per campaign
$0.002–$0.10 / 1,000Adversarial variant search
Property-based, metamorphic and mutation tests over layouts, driveways, homes and packages until a round of variants finds nothing new.
Until saturation
same as L2Bounded agents, mixed model families
Draft responses, extract facts from drawings for a person to confirm, find gaps the rule packs missed, explain. Jev labels and routes.
Tens of calls per run
Jev $0.42–$2.70 / 1,0004D replay of one run
One seeded run's event log replayed on the real lot in the existing Mapbox and Three.js stage.
One run at a time
≈ $0 computeBacktest and calibration
Frozen predictions scored against real letters, permits and inspection results. This is the only layer that makes the others trustworthy.
Each real outcome
staff timeCost per campaign, by how much language model sits in each run
Log scale: each gridline is ten times the last.
Planning assumptions from the runtime research (8 inspection stages, list prices checked 2026-10-06), not quotes. Role-play spans Haiku with caching and batching up to a heavy Opus design, and it adds no ground truth: an agent passing a drawing is not evidence that the drawing complies.
Run the example engine thousands of times
The same engine as the simulator above, with the pro team's answer quality drawn at random. Each setting is seeded, so it replays exactly.
Thousands of simulated GC teams
A GC team is a set of parameters (crew size, response lag, rework rate, season, inspection lead time), each with a labelled prior: its evidence grade and the number of real observations behind it. A thousand teams is a thousand draws, labelled SYNTHETIC, never a bid. A thousand LLM "GCs" would be a thousand near-copies of one model's opinion. Calibration comes from real JobTread and quote data once the cost and schedule spine (spec 0058) lands.
Product packages for costs
The cost tool is a wrapper, not an agent. A package or option changes a run's cost only through @basemod/pricing-engine; anything unpriced shows pricing_pending. Staff assumptions are labelled as such. No agent on either team authors a dollar.
Network effects
Every real outcome makes the next project cheaper
Regulation is layered: one state code, then county agencies, then each township. Onboard a county once and its road, health and drain packs serve every township in it. Each real letter, permit and inspection result calibrates the authority that issued it, so the next project there starts with fewer surprises. The flywheel only turns on real outcomes. Until those are logged, the system has data, not network effects.
Onboard once, reuse many
Pick where the next lot is. The map shows which rule packs carry over and which places need new intake.
© Mapbox © OpenStreetMap © Maxar. Boundaries: State of Michigan.
| Asset that compounds | Reused by | Compounds when | Breaks when |
|---|---|---|---|
| Rule packs per authority | Every run in that authority's territory | Rows are current and human-reviewed | The source changes (marked stale) or its licence forbids reuse |
| Outcome records and golden cases | Every backtest and calibration | Each real letter, permit and inspection is logged per deal, names stripped | Outcomes are not captured. Then nothing compounds. |
| Precedent library | The pro team's fix search | The fix was accepted by the real authority | The fix only ever satisfied the simulator |
| Model passports | Every later lot using that registered home | The footprint is registered in land-core and the rule snapshot is unchanged | The model, the drawing set or the rule snapshot changes |
| Authority rosters | Every parcel inside a boundary | Each activation records its evidence | A boundary or a contracted agency changes |
Roadmap
Prove it on one real lot first
Each phase names one outcome, an owner, what it excludes, and the gate to the next. Local proof is reported separately from signed-in live acceptance.
- Do
- Name the person who adjudicates findings, set the Phase 0 spend cap, start logging every real agency outcome per deal with names stripped, approve human-sent requests for the missing inspection lists, and authorize counsel review.
- LangSmith
- Create the account and pick the plan (agents never create accounts), put the API key in your local environment only, and approve the trace policy (D6).

- Outcome
- Staff see a backtested rehearsal of Robinson Lot 15 across four lanes: health, driveway, soil erosion trigger and zoning certificate, plus a building-permit prerequisite check.
- Built as
- A LangGraph.js graph on your Mac with local checkpoints, traced to LangSmith with masking; the real Robinson outcomes become its first dataset.
- Excludes
- A Vercel deploy, Jev, new Supabase tables, Unreal, Municode text, customer surfaces, outward actions.
- Gate
- Frozen predictions scored against the real health and road outcomes; zero banned strings; every check cites section, edition, retrieval date and hash.
- Outcome
- The full journey on a signed-in staff preview for the Robinson lot: contracts, staff-only schema, rule-pack store, referee, ledger, independent first reviews across model families, response drafting, Review Room cards, approval gates.
- Needs
- ADR 0027 accepted (LangGraph on Vercel), the paid Vercel plan confirmed, the Agents rail merged, and the checkpoint schema created by a reviewed migration.

- Outcome
- 1230 N Prospect, Ypsilanti Charter Township: a forward test of the open-space preliminary site plan, frozen before the next real township interaction, plus a deterministic sweep of the lot-count and private-road boundary cases.
- Needs
- Counsel clears Municode-sourced text, or the township supplies its files.
- Outcome
- Seeded Monte Carlo of rounds, clocks and inspection readiness with labelled priors; variant sweeps; a 4D replay of one run; GC populations once the cost and schedule spine lands.
- Outcome
- Calibrated labels where backtests clear the bars, and a repeatable intake procedure for more Michigan places, chosen by deal demand.
- Outcome
- Partner OS exposure only after backtest numbers, counsel's wording review and licensed human gates. It may never happen, and that is an acceptable outcome.
Built on
- LangGraph
- LangSmith
- Vercel AI SDK
- Supabase
- Mapbox
What the agents never do
- Quote a price. Buyer dollars come from the pricing authority or stay pending.
- Decide fit, zoning approval or who qualifies.
- File, send or sign anything. Drafts leave only through a person.
- Use broker-only or licensed data outside its license. Every guardrail, in full.
BaseMod.ai