Skip to content
BaseMod.ai Private preview Run a rehearsal

BaseMod.ai Regulatory Rehearsal

Rehearse the approval before you file.

BaseMod agents play both sides of a land review. A developer team submits the plan. Simulated road, health, zoning, soil-erosion and building reviewers object, cite the rule and recheck every fix, round by round, on the real lot.

Lot 15 in daylight satellite with its Ottawa County parcel lines. The house, drive and septic field are a screening sketch, and the setbacks shown are unverified.© Mapbox © OpenStreetMap. Imagery © Maxar.
Two teams, opposite goalsThe pro team drives weeks and dollars down. The regulator team drives safety, code and ordinance. Neither can trade away the other's goal.
Only evidence closes an objectionA regulator agent cannot withdraw an objection because it was persuaded. The referee admits one of the closure events, or it stays open.
The thousands of runs are codeDeterministic runs cost fractions of a cent each. Agents are called only where judgment is needed.
Reality is the only teacherReal review letters, permits and inspection results calibrate it. Its own agents agreeing never counts.
Nothing here is an approvalResults are always labelled simulated. People confirm facts, and nothing is filed by an agent.

Product

Run a rehearsal

Pick a venue and step through a review the way the engine runs it. Objections pin themselves to Lot 15 as the simulated agency raises them. You answer each one with one of four responses, the referee decides what closes, and the run ends with a result card. This is the Review Room staff would use in Land Explorer, on a daylight satellite plan of the real lot. Every scenario here is an example, not a real run.

How it works

Two teams, one engine, one referee

The pro team on the left proposes and answers. The regulator team on the right challenges, grouped the way Michigan authorities actually work: the township's own review cluster, the bodies that vote, the county and state agencies you apply to separately, and the building department. In the middle sits the machinery that keeps the contest honest. The engine owns every number and every state change. The referee decides which objections and answers count. You sit at the bottom, in the Review Room.

Select any node to open its role card beside it; Esc closes it. Every card is also in the roster tables further down. On a phone, scroll the diagram sideways.

Regulatory Rehearsal system map The pro team of eleven roles led by the Lead Developer orchestrator on the left; the regulator team on the right grouped into the municipal review cluster led by the Planning Director orchestrator, deciding bodies, independent peer authorities and the building department; in the center the project revision, the regulatory vault, the rehearsal engine, the referee, the objection ledger, the model router and the Review Room.
Pro teamRegulator teamMachinerySourcesYou Submission and re-checkObjectionsRetrieval, same pinned snapshotAdmitted to the ledgerEscalated to you

One round

Twelve steps from a saved project to a result card

This is one review round on one venue, such as the county road commission's driveway permit. The same machinery runs for zoning, health, soil erosion and the building department. Only the rule pack and the authority change. Use Next, or press Play.

Step 1 of 12

Sequence of one review round across eight lanes

Thousands of runs

Where the thousands actually go

Your instinct to test a design thousands of times is right. The research says where to spend them. Repeated runs of the same language model are correlated, drift toward the same answers and cannot be replayed by a seed. So the thousands run in code, on seeded engines, and agents work where judgment is needed. Six layers do the work.

L1

Rule packs: the real counter-agents

Edition-pinned, table-valued checks per authority: setbacks, sight distance, isolation distances, frontage, triggers. Each row cites a section number, a paraphrase, a source date and a hash, never code text.

Once per project revision

≈ $0
L2

Seeded Monte Carlo and discrete-event runs

Review rounds, agency clocks, inspection readiness and crews. Every run carries a seed and replays exactly.

Thousands per campaign

$0.002–$0.10 / 1,000
L3

Adversarial variant search

Property-based, metamorphic and mutation tests over layouts, driveways, homes and packages until a round of variants finds nothing new.

Until saturation

same as L2
L4

Bounded agents, mixed model families

Draft responses, extract facts from drawings for a person to confirm, find gaps the rule packs missed, explain. Jev labels and routes.

Tens of calls per run

Jev $0.42–$2.70 / 1,000
L5

4D replay of one run

One seeded run's event log replayed on the real lot in the existing Mapbox and Three.js stage.

One run at a time

≈ $0 compute
L6

Backtest and calibration

Frozen predictions scored against real letters, permits and inspection results. This is the only layer that makes the others trustworthy.

Each real outcome

staff time

Cost per campaign, by how much language model sits in each run

Log scale: each gridline is ten times the last.

Planning assumptions from the runtime research (8 inspection stages, list prices checked 2026-10-06), not quotes. Role-play spans Haiku with caching and batching up to a heavy Opus design, and it adds no ground truth: an agent passing a drawing is not evidence that the drawing complies.

Run the example engine thousands of times

The same engine as the simulator above, with the pro team's answer quality drawn at random. Each setting is seeded, so it replays exactly.

Thousands of simulated GC teams

A GC team is a set of parameters (crew size, response lag, rework rate, season, inspection lead time), each with a labelled prior: its evidence grade and the number of real observations behind it. A thousand teams is a thousand draws, labelled SYNTHETIC, never a bid. A thousand LLM "GCs" would be a thousand near-copies of one model's opinion. Calibration comes from real JobTread and quote data once the cost and schedule spine (spec 0058) lands.

Product packages for costs

The cost tool is a wrapper, not an agent. A package or option changes a run's cost only through @basemod/pricing-engine; anything unpriced shows pricing_pending. Staff assumptions are labelled as such. No agent on either team authors a dollar.

Network effects

Every real outcome makes the next project cheaper

Regulation is layered: one state code, then county agencies, then each township. Onboard a county once and its road, health and drain packs serve every township in it. Each real letter, permit and inspection result calibrates the authority that issued it, so the next project there starts with fewer surprises. The flywheel only turns on real outcomes. Until those are logged, the system has data, not network effects.

The rehearsal flywheel Rehearse a real project, log the real outcome, calibrate that authority, fewer rounds next time, more lots and developers, more places onboarded, and back to rehearsing. Rehearse a real projectbefore money or forms Log the real outcomeletters, permits, inspections Calibrate that authoritypacks, priors, precedents Fewer rounds next timein that county and township More lots, developersPartner OS, later and gated More places onboardedgoverned intake, by deal Turns only on real outcomes Track 0: log every agency outcome per deal, names stripped

Onboard once, reuse many

Pick where the next lot is. The map shows which rule packs carry over and which places need new intake.

Daylight satellite map of lower Michigan from Grand Haven to Ypsilanti with county and township outlines © Mapbox © OpenStreetMap © Maxar. Boundaries: State of Michigan.
Asset that compoundsReused byCompounds whenBreaks when
Rule packs per authorityEvery run in that authority's territoryRows are current and human-reviewedThe source changes (marked stale) or its licence forbids reuse
Outcome records and golden casesEvery backtest and calibrationEach real letter, permit and inspection is logged per deal, names strippedOutcomes are not captured. Then nothing compounds.
Precedent libraryThe pro team's fix searchThe fix was accepted by the real authorityThe fix only ever satisfied the simulator
Model passportsEvery later lot using that registered homeThe footprint is registered in land-core and the rule snapshot is unchangedThe model, the drawing set or the rule snapshot changes
Authority rostersEvery parcel inside a boundaryEach activation records its evidenceA boundary or a contracted agency changes
When this starts compoundingPhase 4 at the earliest. Until then every rehearsal stays staff-only, and outside users (Partner OS) wait for backtest numbers, counsel's wording review and licensed human gates.

Roadmap

Prove it on one real lot first

Each phase names one outcome, an owner, what it excludes, and the gate to the next. Local proof is reported separately from signed-in live acceptance.

Track 0You, no code
Do
Name the person who adjudicates findings, set the Phase 0 spend cap, start logging every real agency outcome per deal with names stripped, approve human-sent requests for the missing inspection lists, and authorize counsel review.
LangSmith
Create the account and pick the plan (agents never create accounts), put the API key in your local environment only, and approve the trace policy (D6).
Lots 15 to 17 at Grand River Estates in daylight satellite
Grand River Estates, Robinson Township. Lot 15 is the Phase 0 lot. Its real health and road outcomes are the first backtest. Imagery © Mapbox © OpenStreetMap © Maxar.
Phase 0Local, about two weeks
Outcome
Staff see a backtested rehearsal of Robinson Lot 15 across four lanes: health, driveway, soil erosion trigger and zoning certificate, plus a building-permit prerequisite check.
Built as
A LangGraph.js graph on your Mac with local checkpoints, traced to LangSmith with masking; the real Robinson outcomes become its first dataset.
Excludes
A Vercel deploy, Jev, new Supabase tables, Unreal, Municode text, customer surfaces, outward actions.
Gate
Frozen predictions scored against the real health and road outcomes; zero banned strings; every check cites section, edition, retrieval date and hash.
Phase 1Agents rail, Vercel preview
Outcome
The full journey on a signed-in staff preview for the Robinson lot: contracts, staff-only schema, rule-pack store, referee, ledger, independent first reviews across model families, response drafting, Review Room cards, approval gates.
Needs
ADR 0027 accepted (LangGraph on Vercel), the paid Vercel plan confirmed, the Agents rail merged, and the checkpoint schema created by a reviewed migration.
The Prospect Road area of Ypsilanti Township in daylight satellite
1230 N Prospect, Ypsilanti Charter Township. The forward test freezes its predictions before the next township interaction. Imagery © Mapbox © OpenStreetMap © Maxar.
Phase 2Second place
Outcome
1230 N Prospect, Ypsilanti Charter Township: a forward test of the open-space preliminary site plan, frozen before the next real township interaction, plus a deterministic sweep of the lot-count and private-road boundary cases.
Needs
Counsel clears Municode-sourced text, or the township supplies its files.
Phase 3Simulation layer
Outcome
Seeded Monte Carlo of rounds, clocks and inspection readiness with labelled priors; variant sweeps; a 4D replay of one run; GC populations once the cost and schedule spine lands.
Phase 4Calibration and intake
Outcome
Calibrated labels where backtests clear the bars, and a repeatable intake procedure for more Michigan places, chosen by deal demand.
Phase 5Gated outside use
Outcome
Partner OS exposure only after backtest numbers, counsel's wording review and licensed human gates. It may never happen, and that is an acceptable outcome.

Built on

  • LangGraph
  • LangSmith
  • Vercel AI SDK
  • Supabase
  • Mapbox

What the agents never do

  • Quote a price. Buyer dollars come from the pricing authority or stay pending.
  • Decide fit, zoning approval or who qualifies.
  • File, send or sign anything. Drafts leave only through a person.
  • Use broker-only or licensed data outside its license. Every guardrail, in full.

BaseMod.ai

Find every objection first. Then a licensed team files for real.