Systems you can run
Three small working systems. Each one shows what happens underneath: which steps stay deterministic, where a model earns its place, and where a person decides.
Deterministic
AI / probabilistic
Human judgment
Synthetic data
Everything runs in your browser on scripted, invented data. There are no live model calls, and every person, company and record here is fictional.
Project 01 Knowledge
Project 02 Operations
Project 03 Workflow
Concept Synthetic data
Knowledge Intelligence
Build note
The problem Answers to policy questions are spread across documents, some of them out of date, and people can't tell which answer to trust.
Riskiest assumption That retrieval finds the right passage often enough for every answer to cite its source.
What it does Real BM25 retrieval over ten synthetic policy passages, computed as you type. The sample questions have scripted answers. Your own questions get an extractive answer quoted from the sources.
Left to rules and people Every claim cites a retrieved passage. Superseded documents are down-weighted and labelled. Missing evidence produces a refusal, not a guess.
What's next Test it on a larger document set with questions whose answers are known, and measure how often it cites the right passage.
Ask the knowledge base Lumen Ledger · 4 docs · 10 passages
Concept Synthetic data
Operational Intelligence
Build note
The problem Work fails between steps, where no single application can see it.
Riskiest assumption That simple, explainable rules catch the failures that matter without burying people in false alarms.
What it does Four detection rules run live over 214 synthetic customer-onboarding records across CRM, delivery and billing. Each flag carries the evidence that triggered it.
Left to rules and people Detection is deterministic. The system recommends and a person decides, and every decision is recorded so each rule can be measured.
What's next Use the recorded decisions to tune or retire the rules that raise too many false alarms.
Exceptions ↻ Re-run detection
37 records require attention
Concept Synthetic data
Intelligent Workflow
Build note
The problem An inbound lead needs research, a score, a CRM update and a reply, and doing that by hand is slow and inconsistent.
Riskiest assumption That a model's classification is reliable enough to help score a lead, and that the system can tell when it isn't.
What it does A lead arrives by webhook. The system normalizes it, researches and classifies it, scores and routes it, matches it in the CRM and drafts the outreach.
Left to rules and people Normalizing, scoring, routing and CRM matching are code. Low model confidence takes AI out of the score, low-fit leads skip outreach, and nothing is sent until a person approves it.
What's next Compare scores with what actually happened to each lead, using the record every run writes.
Incoming webhook ▶ Run workflow