FLOBOTICSLAB Last run finished 14 s ago

Public engineering laboratory · since 2025

We test it before you trust it.

Models, agents and tools, run against real healthcare workflows on synthetic data. Every experiment has a question, a hypothesis written before the first run, a published protocol and raw logs. Failures stay online.

174Experiments published
Register updated
2026-09-22 19:58 CEST
Concluded
142
Active
3
Inconcl.
7
Failed
19
Retracted
3
Running nowStreaming from the harness · illustrative
lab-harness v0.9.3 · 3 active experiments · 2 GPU nodes · 1 unified-memory workstation
FLX-0149
Three concurrent agents on one prior-auth queueGPT-6 · Claude · Qwen 32B local · since 2026-09-18
31 / 50
FLX-0148
Six local models on 256 GB of unified memoryRequisition extraction · 300 documents · since 2026-09-19
212 / 300
FLX-0147
Replicating: 1,000 runs of one denial-appeal agentSecond engineer · different week · since 2026-09-15
684 / 1000
Experiment register12 shown · 174 total

Every experiment, including the ones that did not work.

Each card shows the question, the status and the result in the form that fits the experiment: lanes for battles, run strips for reliability, thresholds for failures.

0149
FLX-0149 · Agent battleRunning

Three concurrent agents on one prior-auth queue: does parallelism beat one fast agent?

Runs31 / 50
GPT-6ClaudeQwen 32B

Hypothesis: three agents in parallel clear a 200-item queue 2x faster than GPT-6 alone.

Started 2026-09-18Est. cost $140
0148
FLX-0148 · Local benchmarkRunning

Six local models on 256 GB of unified memory: how many agents can one workstation run?

Documents212 / 300
Qwen 72BQwen 32BLlama 70B+3

Hypothesis: planner, extractor and verifier fit in memory at once with 4 concurrent agents.

Started 2026-09-19Electricity only
0145
FLX-0145 · Local benchmarkConcluded

Six local models on handwritten requisitions: accuracy vs tokens per second

Result: Qwen 32B at Q8 gives the best accuracy per second. The 72B model adds 1.8 points at 2.6x the time.

2026-09-05n = 400 docs$0 API cost
0142
FLX-0142 · Agent battleConcluded

GPT-6 vs Claude vs local Qwen: which agent can automate this workflow fastest?

Result: GPT-6 fastest by median. Claude most reliable. Qwen local 13x cheaper per run.

2026-09-12n = 50 × 3$58.40
0140
FLX-0140 · ReliabilityConcluded

1,000 runs of one denial-appeal agent: where do the failures cluster?

Each square is one run · red = failed

Result: 97.3% success. 21 of 27 failures hit 3 payer portals during their 06:00 to 07:00 maintenance window.

2026-08-29n = 1,000$212
0138
FLX-0138 · AI vs humanConcluded

Four billing specialists vs one agent: 200 denial appeals, blind-reviewed

Result: the agent was 4.1x faster; specialists scored 6 points higher. Agent draft plus specialist review scored highest.

2026-08-22n = 200Blind review, 2 judges
0136
FLX-0136 · Cost comparisonConcluded

Cost per 1,000 claim status checks: human vs RPA vs agent vs direct API

Result: use the API where one exists. Agents beat RPA only on portals that change more than once a quarter.

2026-08-15n = 4 × 1,000Log scale
0127
FLX-0127 · Weird ideaConcluded

Can an agent write its own RPA bot? We let it try on 12 workflows.

Worked 7Fixed 4Unsafe 1

Result: useful for scaffolding. The unsafe bot led to a new guardrail: no submit action without a validation step.

2026-07-18n = 12$31
0133
FLX-0133 · Tool testInconclusive

Does a vision model read fax cover sheets better than OCR plus rules?

Usable sample n = 18Required n = 50

Why inconclusive: 32 of 50 samples excluded for image quality. Re-running with a new synthetic set.

2026-08-08n = 18 of 50
0131
FLX-0131 · Voice agentFailed

Can a fully local voice agent handle appointment rescheduling calls?

Threshold 1.0 sMeasured 2.8 s

What we keep: the call-flow state machine, now in production with a hosted model.

2026-08-01n = 40 calls$12
0129
FLX-0129 · Agent battleRetracted

Browser agents on payer portals, round 1

Reported completion: 98%
Cause: the sandbox portal cached answers between runs, so later runs were easier.
Replaced by FLX-0142.

Why it stays online: so anyone who cited it can see what changed.

Retracted 2026-08-30
0147
FLX-0147 · ReliabilityReplicating

Same agent, same task, different week: does FLX-0140 hold?

684 of 1,000 runs so far

So far: 97.1% success, failures again clustered in portal maintenance windows.

Since 2026-09-15Second engineer
Failure wall19 failed · 7 inconclusive · 3 retracted · all online

What did not work, and what we kept from it.

A failed experiment is a result. Most of them leave behind a component, a guardrail or a threshold that now protects a production system.

FailedFLX-0131 · 2026-08-01

Can a fully local voice agent handle appointment rescheduling calls?

Threshold 1.0 s2.8 s per turn

Callers talked over the agent in 31 of 40 test calls.

What we keep

The call-flow state machine. It now runs in production with a hosted model and passes the same test at 0.7 s.

FailedFLX-0120 · 2026-06-20

Can an agent navigate a payer phone tree to check claim status?

9 of 40 callsThreshold 36 of 40

Phone-tree speech recognition misheard member IDs in most calls.

What we keep

A keypad-only path for two payers whose phone trees accept tone input. Used in production since July.

InconclusiveFLX-0133 · 2026-08-08

Does a vision model read fax cover sheets better than OCR plus rules?

Only 18 of 50 samples were usable after image-quality exclusions. Too small to call.

Next

FLX-0151 is proposed with a new synthetic set generated at real fax resolution.

Proposed
RetractedFLX-0129 · retracted 2026-08-30

Browser agents on payer portals, round 1

The sandbox portal cached answers between runs, so later runs were easier. The reported 98% was not real.

What changed

The harness now resets the sandbox between every run and checks for cached state. FLX-0142 replaced this experiment.

MethodsSame protocol for every experiment

How an experiment runs.

The hypothesis and the thresholds are written down before the first run. That is the single rule that keeps a lab honest.

01Question

Plain language, tied to a real workflow a client runs.

02Hypothesis

And pass thresholds, written before the first run.

03Protocol

Dataset, environment and harness version frozen.

04Runs

At least 50 per contender, in random order.

05Replicate

A second run, another week, another engineer.

06Publish

Result, raw logs and harness. Failures too.

Production candidate thresholdMinimum
End-to-end completion≥ 95%
Field-level accuracy≥ 98%
p90 time vs median≤ 1.6x
Unsafe actions (submit without validation, wrong patient)0
Cost per case vs human baseline< 25%
Replication difference≤ 2 points
reproduce.sh · bash
# Reproduce any Lab experiment on your own hardware
git clone https://github.com/flobotics/lab-harness
cd lab-harness && make env
flx run FLX-0142 --runs 50 --replicate
✓ 50 synthetic referrals loadedsandbox reset between runsest. 2 h 40 min

Lab log

One short email when an experiment concludes, fails or is retracted. Nothing else.