BuiltSynthetic data

Business
Diagnostic

Finding where an operation stops matching the process everyone thinks is happening.

A method for any operation that records events. It reconstructs what actually happened from the records a business already keeps, then reports only what the evidence supports. Its first implementation runs on dealership CRM data.

Tue 9:39–9:40Lead received, auto-reply sent, assigned 99.3 hoursof business time with no one touching it

The short version.

Both gears turn. The trouble is where they meet.

Most operational problems aren't obvious failures. The software works, the automation runs, the report gets generated, and people follow the process. Something still gets lost between them.

Those gaps are hard to see because they sit between parts that each look fine on their own. The Diagnostic tests one question: can the event data a business already records expose those gaps reliably enough to produce a finding a manager can check?

It doesn't report what the process diagram says should happen. It reports what happened, the evidence behind it, and what the data can't establish.

Four ways work stops.

Every finding names one of four causes. None of them is specific to cars: any operation with handoffs has all four.

Waiting

Something sat in one state longer than the standard allows.

A lead unworked past the response window.

Unowned

Responsibility existed in theory, but no one owns it in the record.

A lead assigned to someone who left.

Invisible

The process kept going, but an exception dropped out of sight of the people meant to act on it.

A lead closed as lost that nobody ever spoke to.

Duplicate

Several records or paths stand for what should be one thing.

The same customer arriving from two sites.

Follow one lead.

One lead from a synthetic dealership export, shown exactly as each stage of the pipeline produced it. The customer details are invented. It plays on its own until you pick a stage.

01 · Export

Start with what the system already recorded.

The store exports its CRM's lead activity: one row per event. Nothing new is installed and no new source of truth is invented. This lead came in from Cars.com at 9:39 on a Tuesday morning.

CRM export · 3 rows for lead L100329raw
LeadIDCustomerPhoneEmailSourceAssignedToActivityTypeTimestamp
L100329Gregory Tucker817.488.4795susan11@example.orgCars.comB. WalshNew Internet Lead09/15/2026 09:39 AM
L100329Gregory Tucker817.488.4795susan11@example.orgCars.comB. WalshAuto-Reply09/15/2026 09:39 AM
L100329Gregory Tucker817.488.4795susan11@example.orgCars.comB. WalshLead Assigned09/15/2026 09:40 AM

The full synthetic export: 4,227 activity rows across 703 leads, with duplicates, missing timestamps, departed owners and silent losses planted on purpose so the pipeline can be checked against a known answer.

02 · Anonymize

Know it happened without knowing who.

Customer fields are replaced with a salted hash before any analysis, using a salt set for each engagement. The same customer always gets the same hash within a run, so duplicates can still be found. An entity-recognition model scrubs names and numbers out of free-text notes, and if it can't run, the stage stops instead of passing notes through.

Anonymized · same 3 rowsclean
LeadIDCustomerSourceAssignedToActivityTypeTimestamp
L100329Gregory Tucker · 817… · susan11@…
CUST-1A95366377
Cars.comB. WalshNew Internet Lead09/15/2026 09:39 AM
L100329CUST-1A95366377Cars.comB. WalshAuto-Reply09/15/2026 09:39 AM
L100329CUST-1A95366377Cars.comB. WalshLead Assigned09/15/2026 09:40 AM

The redaction log records counts only, never the values it removed.

03 · Normalize

One vocabulary for every system.

Every CRM names the same events differently. A mapping file translates each raw label into a common event, and the rows become a clean event log: what happened, when, to which case, and by whom.

A label the map doesn't know is refused, not guessed. In this run 3 rows had unmapped labels and 137 had unusable timestamps. All 140 went to a data-quality list instead of disappearing.

Event lognormalized
case_idactivitytimestampresource
L100329New Internet Lead → Lead Received2026-09-15 09:39B. Walsh
L100329Auto-Reply → Auto Response2026-09-15 09:39B. Walsh
L100329Lead Assigned → Lead Assigned2026-09-15 09:40B. Walsh

Changing CRMs means writing a new mapping file. The analysis code doesn't change.

04 · Reconstruct

The sequence is the unit.

A lead being created isn't interesting by itself. Neither is an assignment. The information is in the relationship between events, and in what never followed them.

Time is measured in the store's business hours, so a lead that arrives after close starts aging at the next opening.

Lead L100329 · as of Wed Sep 23, 9:00 PM
  1. Tue 9:39Lead receivedCars.com · Kicks SV
  2. Tue 9:39Auto response sentsystem activity, doesn't count as work
  3. Tue 9:40Assigned to B. Walshadmin activity, doesn't count as work
  4. 8 daysNothing99.3 business hours with no call, text, email or appointment
05 · Analyze

Test it against the store's own rules.

Every rule belongs to the dealer: store hours, what counts as working a lead, how long a lead may wait. They live in a versioned file, and any rule still on a suggested value is printed in the report.

Two checks fail for this lead. The pipeline only measures. Naming the cause is an analyst's call, recorded with its evidence.

Checks · lead ownership rules v0.1
  • noWas it worked within 4 business hours?No human attempt at all. Auto responses and assignments are excluded on purpose, so a system message can't make a lead look handled.
  • noIs the owner on the active staff list?B. Walsh is on the roster as inactive. The lead belongs to someone who no longer works there.
  • likely causesWaiting · UnownedThe cause narrows the investigation. It doesn't decide the fix.
06 · Report

Something a person can act on tomorrow.

The lead lands on the next morning's exception list, routed to the person who can fix it. The list carries lead IDs, not customer details. The store looks the lead up in its own CRM.

Across the whole store, patterns like this one become findings, which have to pass the finding contract below before they reach a report.

Morning list · 32 exceptions from the last 14 days
L100329 · Cars.com · Kicks SV99.3 business hrs

Untouched 99.3 business hours. Owner not on active staff list.

→ Sales manager: assign an owner

In this synthetic run the classifier recovered every planted case it was built to find. A test checks that on every change.

Why the stages are separate.

Each boundary exists so one kind of change can't break another.

Source systemAnalysis

A new system

Changing how a CRM names an activity means editing a mapping file, not rewriting the analysis.

RulesEvidence

A new rule

Changing an operating standard is a versioned config change. It never rewrites the historical evidence.

FindingsReport

A new report

The report layer can't make a weak finding stronger by writing more confident prose.

Fail closed.

A finding is a structured record, not a paragraph, and missing evidence makes it weaker, never louder. Switch the evidence on and off to see how the rules respond. This is a simplified version of the schema the pipeline enforces.

When it finds something

Result

Headline: “Leads sit unworked after hours.” The wording doesn't change. Only the confidence label does.

When it finds nothing

Result for a clean run

A clean result is only reported as clean when coverage is established. Otherwise the report says there wasn't enough evidence to know.

When things go wrong.

The happy path is the easy part. These are the cases the code actually handles.

An export never arrives, or has a gap

A preflight check runs before any analysis and answers READY or NOT READY. Anything short of a pass, including a check it couldn't run, means not ready. Known missing periods mark the result incomplete, so a clean result is never reported over a gap.

The CRM renames an activity

An unknown label is refused, not guessed, and logged as a data-quality item. The preflight blocks until every raw label has a meaning in the mapping file.

Two records look like the same lead

Matching leads are grouped into one opportunity, so the same customer isn't counted twice. Known gap: matching is exact today (same name and phone). Real exports format phones differently, and fuzzy matching isn't built yet.

The owner is missing or has left

Every owner is checked against the store's staff list. A blank owner or someone no longer active makes the lead unowned, and it's routed to the sales manager to reassign.

Timestamps are missing or impossible

A lead with no usable received time goes to the data-quality list, not the findings. Events dated before the lead arrived, or in the future, are flagged the same way. Nothing is silently dropped.

The same evidence could support two explanations

Findings from different methods are only merged when a person declares the link, and only within the same workflow. Nothing is merged because it looks similar.

Someone edits a finding to sound stronger

Analyst notes can add evidence but can never create a finding or change a pipeline fact. If one tries, the run fails and writes nothing. Every sentence in the report is traced back to its source, and a headline can't contain numbers, so it can't disagree with the counts.

Real customer data ends up in the code repository

It doesn't go there. The repository and its automated tests only ever see synthetic data. Real exports stay in the store's own isolated environment under a data agreement.

Testing the boring parts.

1,000automated tests, all passing. The number matters less than what they protect.

  • PrivacyPlanted personal details must be removed from notes, while staff names, vehicles and dates survive.
  • Known answersEvery planted case in the synthetic export (duplicates, orphans, silent losses, bad data) must be recovered.
  • No inflated claimsThe validator re-runs the rules on the recorded counts, so a document can't claim a finding its counts don't support.
  • Nothing changes by accidentEarlier report versions are pinned byte for byte, so a new version can't quietly alter an old one.

How it fits any business.

Dealerships are the first implementation, not the definition. Every business has its own systems and its own standards, so every engagement starts by mapping one and declaring the other. The rest of the method is the same everywhere.

The same for every business

  • The event model · what happened, when, to what, by whom
  • The four causes
  • The finding contract and fail-closed validation
  • Anonymize, normalize, reconstruct, analyze, report
  • The report checks

Set up for each business

  • The source mapping · translates that business's system labels into the event model
  • The rules being tested · the business's own standards, declared before the data is read

Every engagement includes both steps, dealerships too. The mappings and rules written so far are for a dealership CRM.

A service company's ticket queue

Tickets that wait past the response standard, or sit with a technician who's off the schedule.

A clinic's referral intake

Referrals received but never scheduled, or entered twice from fax and portal.

A B2B inbound sales desk

Demo requests assigned to a rep who left, with nobody noticing them go quiet.

Examples of where the same method would apply. None of them has been run.

Status.

Engineering evidence isn't the same as proof in a live operation. Both steps are listed.

  1. Runs end to end on synthetic data

    Export to morning list and finding report, checked against a dataset where the right answer is known.

Build note.

The problem, the assumption, the system, and what comes next.

The problem
Operations lose work in the handoffs between working tools and capable people: work nobody owns, waits nobody sees, and exceptions that fall between systems.
Riskiest assumption
That a business's own exports hold enough evidence to show where those handoffs fail, before any new software goes in.
What it does
Anonymizes an export, turns it into an event log, reconstructs each case, tests it against the business's declared rules, and reports findings that pass a strict contract.
Left to rules and people
Every threshold belongs to the business. Analysts can add evidence but can't change a pipeline fact. A person decides what to do about every finding.
What's next
The first run on a real store's export, which needs a mapping for that store's CRM.