Act
Bob listens and answers on the box. Every turn — what it heard, what it said — is logged as an episode step.
Bob is my live RL and ROS laboratory: a local assistant learning from human correction today, and a robot learning to perceive, map and navigate next. Every milestone is published with the evidence — or the failure — that earned it.
Waiting for word from the lab.
Make the map mean something. The lidar gate is paid for — /scan has been
publishing on the robot since 18 August. Since 19 August the odometry gate is paid too: the wheels report
64 times a second and the robot has driven its first metres. What is not paid for is the number — no
driven circuit has come back to its start with the error measured in centimetres.
Wheel odometry now reaches the mapper at 64 Hz — encoder packets stream off the base every 15 ms — and the robot has driven while mapping: commanded 45.8°, measured 45.3°. What remains is the claim itself: drive a circuit, return to the start, publish the gap in centimetres. A map nobody has checked against a return is not yet a localization result.
/scan is live1,080 bins a revolution, 987 with a return, 360.0° at 0.334°, out to 9.34 m.
Scans arrive at 22. The base turns about 53° between two odometry samples, and a straight line drawn through that is how a map smears on every turn.
A driven goal, behind a bumper switch that can stop it in hardware.
A map built from a stationary scanner tests the scanner, not the mapping.
The stand-in node is deleted. The base streams encoders every 15 ms; /odom publishes at 64 Hz. Scanner, mount and mapper unchanged.
Drive a circuit of one floor and measure where the start wall and the end wall land.
Both. A map that looks right in a screenshot is not a measurement.
Record the odometry rate, the turn rate and the divergence, and say which one was wrong.
Navigation stays locked until a bumper signal can stop the robot in hardware.
Model, firmware and health returned successfully.
The blocker was never the jumper. The launch file asked for a scan mode this firmware does not have, so the topic existed and never published. In Boost mode the same sensor went from 720 points at 47.5 per cent to 1,080 at 91.4.
The mapper ran and one map was saved — 7.6 × 9.65 m at 5 cm. It proved the stack, not the localization.
The stand-in node is deleted: encoders stream every 15 ms, /odom publishes at 64 Hz. Commanded 45.8°, measured 45.3. 80 m² of one floor mapped while driving; four rosbags hold the runs.
The project crosses four fields. The tabs separate what each one contributes, what is running now and what evidence is still missing.
Human correction becomes a reward signal, then a confirmed example, then a candidate policy change. Today the data and evaluation layers exist; policy optimization has not run.
32 machine proposals still await a person.
A six-dimensional bandit and 9–22 supervised episodes/hour do not justify policy gradient yet.
A finished run must improve the locked benchmark without regressing held-out cases.
ROS is the nervous system: sensors publish observations, transforms say where they happened, and navigation turns a policy decision into a bounded command.
/scan is publishing987 ranges a revolution at 22 scans/s, since 18 August. Four rosbags since 20 August — runs now outlive their sessions.
The mapper runs on real 64 Hz wheel odometry. The unmeasured loop-closure error is the current gate.
Navigation waits behind localization and a physical bumper signal.
Reasoning is not a personality claim. It is the observable path from a question, through relevant context and checks, to an answer or action that can be graded.
Memory and current sensor state enter the turn without dumping the full corpus.
Code decides measurable facts; a model may propose or critique them.
Wrong tool, wrong context and wrong conclusion are counted separately.
Bob becomes useful by choosing and operating tools: voice, retrieval, vision and eventually ROS. Intelligence here means completing a bounded task with a replayable trace.
Do not run vision for a question that needs only memory.
A tool call is allowed only when its inputs and current hardware are real.
Tool output, latency and result become part of the episode record.
Green means the artifact exists. Amber means the hardware or code exists but the proof does not.
Inputs and scoring exist before optimization.
Only operator-confirmed rows may enter the policy pass.
The live count moves only after a completed training row.
Publishing on the robot since 18 August, 91.4 per cent of bins carrying a return.
Four rosbags written 20 August — 34 minutes, 420,422 messages of /scan, /odom, /tf, /cmd_vel, /bumper. The mapper can be retuned against a drive instead of re-driving the house.
Maps now come from driven runs on real odometry — 80 m² of one floor at 5 cm. No circuit has returned to its start with the gap measured; that is the pass condition.
The robot senses and does not yet move under its own map. Each component stays attached to the limitation that matters — and a bumper switch, the one part that could stop it in hardware, is still not wired.
/scanHealth OK, firmware 1.25. The motor overspeeds at 21.9 rev/s against a nominal 10.
Moved off the desk onto the robot. Useful at distance, blind at bumper range.
Scanner, camera, mapper and console on one 4 GB board, load about 4.4.
A node publishes the identity transform on purpose. Until the wheels report, no map survives a turn.
The training ledger for the body, same rules as the conversation lane: every run is a row, the failures ship with the wins, and a grade only appears after a human confirms it.
| DATE | DRIVER | DIST | BUMPS | OUTCOME | GRADE |
|---|---|---|---|---|---|
| 2026-08-20 21:12 | explore+bump-memory | 14.00 m | 20 | Biggest map yet: 87.5 m2, incl. a 3.3 m ten-goal ZERO-BUMP round - but kept re-electing the same wall corner until boxed | — |
| 2026-08-20 13:55 | beacon-route | 1.50 m | 9 | aborted - operator requested shutdown for updates | — |
| 2026-08-20 13:40 | beacon-route | 1.30 m | 1 | FALSE ARRIVAL - phantom blue: kitchen LED reflection hue 134-140 | — |
| 2026-08-20 13:25 | sweep | 0.00 m | 0 | KITCHEN CONFIRMED: oven 4.9m + microwave 2.7m; waypoint banked | — |
| 2026-08-20 13:05 | beacon-route | 0.70 m | 16 | couch/fridge route failed - started inside pocket | — |
| 2026-08-20 12:55 | explore+camfusion | 2.60 m | 117 | jailed by own glove keep-out disc at reset area | — |
| 2026-08-20 11:52 | patrol+explore | 12.70 m | 20 | wedged in reset pocket; extraction into wall alcove | — |
Ungraded rows are honest rows — the dash means no human has judged that run yet. Full sensor recordings of every episode exist as rosbags on the robot; what is published here is the ledger, never the house.
Bob listens and answers on the box. Every turn — what it heard, what it said — is logged as an episode step.
A reply is graded +1 or −1 with a failure tag, and that grade is the reward signal — no rating scale, no crowdworkers. The count below was entered by an operator over a backlog; the queue waiting on a person is larger than it.
Each bad answer is linked to what should have been said instead. Bad → better becomes a preference pair.
Pairs and grades are banked for the next policy update. Nothing trained yet — the badge stays amber until a run actually consumes them, and the count below is read from the machine, not typed here.
Stage four says banked because it is banked: grades and pairs are accumulating, and no run has consumed them yet. When the first one does, that badge flips — and this page will show it the same way it shows everything else.
Training since January 2026. The results are internal and they vary; what is public is the work itself, and every item below is running, not planned.
There is no chart of a rising line here. Eleven grades do not make a curve, and a graph that steps up and then falls over is a picture of a small sample, not of progress. When there is enough signal to plot honestly, it will be plotted.
Ten of the first eleven grades were failures — grading started with the backlog of known mistakes. Published as-is, because a learning curve that starts at the top isn’t one.
Corrected 14 August 2026. This row used to read “graded by hand”. It was not: all eleven were written through an operator account in a single twenty-three millisecond pass over a transcript backlog, not typed into the console one at a time. The console is built for people and has barely been used — the last login of any kind was 31 July. Meanwhile the grades that did come from real conversation are the row below them, and not one has ever been ruled on. The number was not inflated, but it was described as something it was not, and this page does not get to do that.
Every grade carries a reason, not just a sign. The tag is what makes a failure countable — and what tells the nightly run which behaviour to write a rule against.
823 answers for 0.019 dollars of electricity, drawn from measured wattage rather than a spec sheet. The comparison is the same token volume priced at frontier per-token API rates — not a like-for-like capability claim: a model on a desk is not a frontier model. What it is like-for-like on is this workload, in this building, with nothing leaving it.
/scan: 1,080 bins a revolution, 987 of them carrying a range, a full 360.0° at 0.334°, farthest return 9.34 m in a room the sensor is rated to 12 in. First light 18 August 2026. The number we are not proud of ships with the rest: it spins at 21.9 rev/s against a nominal 10, because the motor-speed line is tied to a 5 V rail instead of being driven at 3.3 V logic.18 August 2026. This paragraph used to refuse two claims: that any of this was scanning LiDAR, and that any of it was stereo depth. Both refusals are now spent — the scanner publishes ranges, the stereo camera publishes depth, and the figures above were read off both. The date is on it, as promised. What is still refused is the one that matters most: this does not have a map it can trust. The mapper runs, and one map has been saved, but the odometry is real now — 64 Hz off the wheel encoders since 19 August — and four rosbags hold the drives that built it. What has never happened is the check: a driven circuit returning to its start with the error published in centimetres. Mapping and navigation stay locked until it does.
Bob is not one computer, and since August he is not two. A graphics box next to the room does the hearing and the remembering; a small desktop machine does all the thinking; and a single-board computer on the robot carries the scanner and the camera and runs their drivers itself. The rule between the first two has not changed: if nobody else is using the graphics card, Bob may. The moment a real job wants it, he moves both his eyes and his brain to the other machine and gets out of the way. The third one is not negotiable in the same way — range data that has to cross wifi before anyone believes it is range data nobody should believe.
Hearing, speaker identity and retrieval run on the GPU, next to the room: a microphone array, transcription, and a voice back out.
A small model takes every ordinary turn. A larger one wakes only when someone asks for depth — and evicts the small one when it does, because both will not fit at once.
A single-board computer on the robot runs the scanner, the stereo camera, the mapper and its own console — on four gigabytes, because everything it does is sensing rather than thinking.
That rule was in the code for weeks and quietly broken by how it was measured. It asked whether the card was full — and Bob's own models filled it, so the card looked busy to Bob and he exiled his own eyes to the slower machine all day. It now asks whether someone else is on it, which is what the rule always meant.
Nothing in this loop is a simulation. The reward comes from people talking to him, the confirmation comes from a person, and the update happens at half three in the morning on a machine in the lab.
What a person says next grades the reply before it. A correction is a negative; carrying on is not. Nobody is asked to rate anything out of five, because nobody does that honestly for long.
The automatic grade is a proposal with a status of its own. It becomes training data only when a human confirms it — and right now 32 proposals are sitting unconfirmed, because nobody has opened the console since 31 July. The gate works. It is the queue behind it that is the problem.
When two trainers split on the same reply it is meant to be shown as a split, not smoothed into a mean. It has never happened. There has only ever been one grader, so the most informative rows in the table are rows that do not exist yet.
Overnight the confirmed rewards distil into behaviour rules that load with him the next day. Every version is kept, so a trainer can see that their grade moved a rule.
His character is a document a person wrote; the policy is what grading distilled. The policy loads last, so on a contradiction it wins. The console now names those pairs instead of letting it happen in the dark.
Found on 14 August, published the same day. The nightly distillation does not read the confirmed grades — it reads the raw reward stream, unconfirmed rows and all. On 8 August a bad wake word opened a listening window that a live football broadcast then held open for seven minutes, and the commentary was transcribed as if it were a person reacting to Bob. Two of those rows scored positive. So the rules count in the strip at the top of this page was distilled from a pool that is roughly three-quarters unconfirmed machine grades, some of which are a television. The fix is a voice check that was installed weeks ago and never wired to anything. Until it is, that number is the least trustworthy figure here, and it is labelled as such rather than quietly removed.
What this page will not show you: the rules themselves, the transcripts they came from, or anything he has remembered about the people who live with him. The shape of the learning is the interesting part, and it is public. What it learned about a family is not.
This is a curriculum built around a real machine, not a claim that the curriculum is complete. The discipline transfers from my production work: build the deterministic reference, lock the benchmark, instrument the run and make a failure reproducible. The new material is policy optimization and the ROS stack that turns a decision into motion.
Deterministic verifiers, seeded fault injection, held-out evaluation, production model serving and a repeatable fine-tune/evaluate/merge/serve pipeline. I know how to make a result falsifiable and how to find the harness bug that makes a model result look mysterious.
Policy search beyond supervised fine-tuning; ROS messages, frames and clocks; sensor bring-up; SLAM; navigation; and the gap between a policy that scores in a frozen environment and one that can move safely in a room. I have not done RLHF or policy-gradient training. That is the gap this lab is designed to close.
trial_eval.py already evaluates frozen map records. The
first lesson is to lock inputs, seed and scoring before tuning anything,
so a better score means a better policy rather than a moving test.
The current controller is a six-dimensional bandit, with roughly 9–22 supervised episodes per hour. At that rate, black-box search is testable in hundreds of evaluations; policy gradient would demand evidence this setup cannot honestly supply yet.
ROS publishes /scan: 987 ranges a revolution, 22 times a
second. Half of this gate is still owed — no rosbag has been
written, so the stream is live and nothing is preserved.
Mapping passes when a driven circuit returns to its start and loop closure error is reported in centimetres. The mapper now runs on real 64 Hz wheel odometry and the robot has driven its first metres; the circuit and its error are still owed. Running the packages is not the experiment.
The order matters: benchmark before optimization, observation before mapping, localization before autonomy. A bumper switch will add a real safety signal; a television in the camera frame does not become a training environment just because the robot can read it.
The blocker was never the wiring. The launch file asked for a scan mode this firmware does not offer, so the topic existed and never published — which looks exactly like a dead sensor. In Boost mode the same scanner went from 720 points at 47.5 per cent valid to 1,080 at 91.4. One line of configuration was worth more than every hardware fix attempted before it.
Two mono sensors 7.5 cm apart, with the disparity computed in fixed-function silicon on the camera's own chip. Accurate to about one per cent at six metres — and unable to see anything closer than 450 mm, which is bumper height. It now rides on the robot beside the scanner, which is the arrangement that makes the blind spot a scanner problem rather than a camera problem.
The board that arrived is a 4 GB Pi 5, not the 8 GB one it was bought as, which matters because the mapper and the camera's object detection are being asked to share it. All four processes run today at about 760 MB and load 4.4. The robot needs onboard compute because mapping wants scan and wheel odometry on one clock.
A sensor mast, a lidar cradle, a guard cage and a camera bracket, modelled to the sensors' real dimensions and printed here. One tripod thread carries the same part through bench, bin and rover, which is the only reason three phases do not need three brackets.
The first maps built by a moving robot landed 19 August: 80.2 m² of one floor at 5 cm a cell, on wheel odometry streaming at 64 Hz. Commanded 45.8°, the encoders measured 45.3. That proves the robot can map while it moves; it does not prove the map is right. The number that will count is whether a driven circuit of one floor has the start wall and the end wall land in the same place.
The honest summary of the robot is that it now has good eyes and the beginning of a memory. It has driven tens of metres, found a person, and read a logo off a television — then failed three times running to reach the kitchen, because it steered toward whatever looked clear and had no map of where it had already been. Since 19 August it builds the map while it moves: the two-a-second wheel reports that blocked this are gone, replaced by an encoder stream at 64 a second. The number still owed is the loop-closure error of a driven circuit — that measurement is the whole of the current experiment.
The training console puts ROS beside reinforcement learning deliberately. The lesson is the bridge: how a correction becomes data, how data becomes a policy candidate, and how a candidate is evaluated before it is allowed to move hardware. I am building that bridge in public because it is the part I am here to learn.
Most of the speed was not bought. It was found.
The eyes moved back onto the free graphics card, and the model was told it is looking at a person rather than at an object being held up.
The frame only reaches the vision model when the question is about seeing. Most turns are not, and now cost nothing.
His character alone filled 86 per cent of the old window, leaving almost nothing for the conversation itself.
Graders were marking down sentences that had been guillotined by a setting rather than by Bob running out of things to say.
Told the wrong name for his own hardware, the old model agreed and wrote it into memory. The new one says no, and says why.
The look was the interesting one. It had been reasoning in circles hunting for an object that was not in the frame, returning nothing at all, and being asked a second time — so the cost was doubled by a sentence of prompt rather than by anything about the hardware.
The one that is not solved: he still re-reads his entire character before writing a word, and on an ordinary turn that is most of the wait. Streaming the answer out sentence by sentence would take the first word from about thirty seconds to under ten without changing a single model. It is written down, it is not built, and this page will say so until it is.
One person, one lab, and a system that has been in training since January 2026. This page used to say two. A second trainer account was created on 31 July and has never been logged into, so the second name came off rather than stay on a page that counts things for a living.
Runs the project, and is the whole of it: writes the code, wires the hardware, grades the replies, and decides what the system is allowed to become.
The console was built for more than one person, and the disagreement machinery behind it has never run because a second person has never used it. The seat is real, the account exists, and it is empty.
Not the company and not the product — the result. What a correctable system looks like when the person correcting it is the one who lives with it.
This page changes as the work does, and the work is not finished. The next thing we are building is the part that makes him answer while he is still thinking, instead of after. After that, the same loop pointed at movement.
If you want to watch it happen rather than read about it later, ask for a walkthrough — or an account on the training console and grade him yourself.
On method, plainly: we train on GPUs we own, using a mix of open-source and commercial models. We do not publish which, because the models are the replaceable part. Results are internal and they vary. The part worth showing is the loop, and that is on this page.
A division of CareCompile. Reinforcement learning and robotics, trained on hardware we own, in the place it runs.