Contents
The first Hynebot prototype in March 2021: a tall birch plywood tower on a plywood base, with a screen and depth camera at the top and a blue robot arm on a carriage halfway up. Hynebot V4 close up: a plywood column with a round speaker, a small camera and a tilted screen showing the J. Hyneman Center logo. Hynebot V4 seen from above, driving through a university lobby between sofas and people walking past.
Left: the first prototype, March 2021. Project archive, photographer unknown. Centre: V4 in 2026. Photograph Teemu Leinonen, LUT University. Right: V4 driving through the LUT lobby, 2026. Photograph Marko Kasurinen.

JHC Platform Series · no. 2 · design record · five generations of one robot, mistakes included

Building a robot for an open workshop

What a university workshop actually needs from a robot, what five generations of one telepresence robot taught us, how the newest of them is learning to answer in Finnish on a single edge computer, and what is still undecided. The principles are meant to travel. The numbers are ours, measured, dated and sourced.

From JHC / Hynebot, V0–V5 Period 2019–2026 Series JHC Platform Series, no. 2 Revision 1.0 · 9 Oct 2026 · first edition

Why this robot exists

Hynebot is a platform project at the J. Hyneman Center (JHC), the open prototyping workshop of LUT University in Lappeenranta. It started in 2019 from a simple observation: student projects in the workshop stall when the person who could help them is not in the room. A robot that lets an instructor be present from somewhere else, look at the part on the bench and point at the problem, seemed like a better answer than another video call. The first version was built by student course teams in 2019–2020, while a partner workshop in San Francisco built a sister robot, and the idea was presented at the SEFI engineering education conference in 2020. SEFIJUV

That is still the first reason, and for most of the robot's life it has been the only one: from 2019 to 2026, Hynebot has been a telepresence device, a machine that a person somewhere else drives, sees and speaks through. The second reason is the same as for every platform project at the J. Hyneman Center. The robot gives students a practical project to put beside the theory: course work, summer jobs and master's theses with a deadline, a budget, parts that arrive late and consequences when a number turns out to be wrong. Each generation has been built, rebuilt or extended by a new group of students and staff (course teams, three master's theses, summer projects). That makes the robot a teaching object in its own right, and a steady source of project and thesis topics.

A third purpose arrived in 2026. The newest version, V5, is being built to remain a telepresence robot and to become an assistant as well: a machine that can answer a question about the laser cutter in Finnish or English, from the laser cutter's own manual, without sending anything to the cloud, and that can eventually drive to where the question is. That adds a degree of autonomy none of the earlier versions had.

These purposes pull the design in different directions, and most of the mistakes recorded here come from letting one of them quietly take over. The version in use today, V4, drives around LUT's buildings as a telepresence robot. V5 has a working brain on the desk and no body yet.

This record is written so that whoever picks up the next version can see why things are the way they are. It is also written for anyone building a robot for a workshop, a lab or a classroom, because most of what we learned is not specific to our parts.

00

How to read this

General guidance and one project's case, kept apart on purpose. Only the first travels to your project unchanged.

This is a design record, not a build manual. Robotics is too wide a field for one project to write a handbook of it, so the book does something narrower: it follows one robot through five generations and records, for each decision, what was decided, why, and what we know now that we did not know then.

How the book is organised

Four parts. Part I is deciding what the robot is for, and the four earlier robots that taught us the answer. Part II is the brain: the speech pipeline, the language model, and the measurements that had to be made before V5 could be built. Part III is the body, largely still in design, and the safety and privacy work that has to come before it. Part IV is method and what comes next. A checklist, the open items and the list of references close the book. V0 to V5 name the robot's generations throughout; chapter 02 opens with a timeline and a table of all of them.

Two kinds of claim

The running text is general; it should hold for any workshop robot. Where a principle meets a concrete decision or number from Hynebot, it is set apart in a box, so the two kinds of claim cannot be confused. There are four kinds of box:

The four box types
BoxWhat it holds
Our caseAn actual decision, measurement or mistake from Hynebot, naming the version, the date and the source. Evidence for the general text, not a recommendation to copy.
RuleThe transferable rule, the thing to carry to a project with entirely different parts.
In plain termsA term the argument depends on, explained where it first carries weight. Skip it if you know the term.
WarningSomething that goes wrong, or went wrong for us. Read these even if you skip everything else.

How to read the numbers

Every figure carries a small reference key that points to its entry in the References, grouped by kind: our own measurements, our own documents and code, theses and conference papers, and published literature. The text stands on its own; the key is there for when a figure matters to your own decision. A number without a key is marked as an estimate. Language-model quality scores are on a 0–5 scale unless stated otherwise. Where the project's own sources disagree, the disagreement is listed under Open items rather than resolved silently.

In plain terms

A word on the word “robot”

Strictly, a robot senses its surroundings and acts on them with some degree of autonomy. A telepresence device does neither on its own: a person drives it and decides everything. By that definition V0 to V4 are telepresence devices rather than robots. V5 is the first Hynebot that earns the name, because it will answer and, later, move without an operator. The project has called every version a robot since 2019, and this book keeps that habit. Where the distinction matters, the text says telepresence or autonomy explicitly.

Where this comes from, and what that limits

The book is written out of one project's material: a 2020 conference paper, requirement lists and safety tables, three theses, a build report, source code from four generations, the V5 measurement campaign of 3 360 language-model answers, and safety and privacy memos. The language-model results rest on one computer, one question set and one automated judge; chapter 10 lists their limits. The history of V1 to V3 is reconstructed from documents, and parts of it are thin. And no student has yet used V5. Chapter 11 returns to that, because it is the most important limit in the book. What transfers is the order of decisions and the ratios between options; individual scores and seconds do not.

This is the second book in the JHC Platform Series, open design records of the J. Hyneman Center's long-running platform projects; the first is on the Ukkonen electric race motorcycle. The book is meant to grow with the project: a superseded figure is replaced in place, and the change is recorded in the revision notes.

Part I

Deciding what the robot is for

Before a single component is chosen: which of the three jobs a workshop robot can do is the one you are building for, and what four earlier robots taught us about the cost of wanting all three at once.

01

The purpose comes before the parts

“A robot for the workshop” is three different design problems. Each one gives you a different machine, and each one needs a different test of success.

A race motorcycle has one metric nobody can argue with: lap time. A workshop robot does not, and that is the first thing to deal with, because a project without a metric will be steered by whatever is easiest to improve that week. In our experience that is usually software.

Three kinds of useful

You can build a workshop robot to be useful in three ways, and they lead to three different machines:

Three kinds of useful, three different optima
GoalWho it servesWhat it leads toHow you know it works
Presencesomeone who is not in the roomgood camera and audio, reliable remote driving, secure access, eye-level screenpeople use it instead of a video call
Assistancesomeone at the bench, nowspeech in both languages, correct answers about specific machines, low latency, mobility to where the question isstudents ask it things and act on the answers without getting hurt
Demonstrationstudents and visitors who watch itvisible mechanism, motion, things they could build themselves, extendabilitystudents want to build the next part of it

The three overlap, which is exactly why it is easy to believe you can have all of them. A presence robot needs mobility and a screen. So does an assistant. A demonstration robot needs both of those plus something that moves in an interesting way. So the requirement list grows, and nothing on it looks unreasonable on its own.

Our case V1 · 2020 requirement list REQJUV

The first requirement list, started in August 2020, asked for a robot that could go through a 90 cm door, fit in a lift with a guardian, climb a 2 cm threshold, grasp and point at things, push buttons, carry its camera at face height with pan and tilt, run for two hours, pack flat for shipping, meet the relevant robot safety standards, and cost under 5 000 dollars. As wishes, it should also lift 1 kg at half a metre, open doors, and act as a laptop stand.

By 13 October 2020, a month in, the cost line already carries the note “Will be exceeded”. It was, by roughly half. Every line on that list was sensible. Together they described a mobile manipulator, and a mobile manipulator is a research platform, not a first prototype.

The viewer cannot tell whether the motion is autonomous

For the demonstration purpose there is a specific trap, and it is worth naming because engineers walk into it on pride. Full autonomous navigation is the hardest and riskiest part of a mobile robot. It is also the part a visitor cannot see. A robot that drives up to you because an operator steered it looks exactly like one that planned its own path. What visitors react to is that it moves, that it speaks their language, and that it comes when called. None of those requires a map of the building.

Rule

Separate mobility from autonomy. Mobility is what the audience sees; autonomy is what the engineer is proud of. Build the first, and put the second behind a decision gate that real use has to open.

“Comes when called” can be done with a marker or person-following and a distance sensor that stops the robot. That is days of work. Free navigation anywhere in the building, with mapping, localisation and path planning, is months, and it is the one workstream that can swallow a small project whole if it is started too early.

The subsystem that grows fastest will try to become the product

The second trap is quieter. In a robot with a language model inside it, the language-model work is the fastest thing to iterate on. It runs on a desk, needs no parts, and produces numbers every day. Left alone it drifts towards a desktop voice assistant, and a desktop voice assistant is one step from a generic chatbot. That teaches students nothing about building machines, and it is not what a workshop is for.

Our case V5 · June 2026 · a correction STATPS

By June 2026 V5 had a working speech pipeline, a chosen model and 2 880 blind-scored answers behind that choice. The plan at that point was to put the brain in a box on a desk first and add mobility “later”. On review we found that plan was the drift described above. A voice assistant in a plywood box is close to a generic AI chat, and it does not show anyone that robots are built here.

The plan was corrected on 15 June: the desk version stays, but as a short step, and the mobile platform moved up to a core goal for the end of 2026, with its component purchases on the critical path. The goal itself had not changed. A robot that moves around the lab and helps the people in it was the aim from the start; it was the AI work that had drifted.

Rule

Write down which purpose leads, and check every quarter which subsystem has actually been getting the hours. The AI is a subsystem of the robot, not the other way round.

What the metric is, for us

For V5 the lead purpose is assistance with mobility: a robot that drives to where a student is working, answers questions about the workshop's machines correctly in Finnish or English, and can be extended by students. Presence remains. V4 serves it today and stays a telepresence robot; V5 is built to serve it as well, as a telepresence robot that can also act on its own. Demonstration follows from the other two if they are done in the open.

The metric that follows is uncomfortable because it cannot be measured on a desk: people who did not build it use it, and what it tells them is right. As of this revision, no student has used V5. Every number in Part II is a measurement of parts of the system under controlled conditions. That is necessary, and it is not the metric.

What we said in 2020, and what changed

The purpose was written down once before, in the JHC team's 2020 conference paper. It is worth quoting how clearly it drew the line. The telepresence robot, it said, does not function like a programmed robot that would require artificial intelligence: “the aim is not to help at some routine tasks, such as safety and lab instructions, answering the same type of questions over and over, monitoring students at work”. It is a camera eye and an arm connected to a microphone, a loudspeaker and a moving platform, all controlled remotely, a tool that wakes up only when someone connects to it. SEFI

V5 is being built to do precisely the routine tasks that paragraph excluded. That is not a contradiction to hide. In 2020 no language model that could run on a robot answered a technical question correctly in Finnish; in 2026, chapter 05 shows, several do. The presence purpose did not go away: V4 serves it, and V5 keeps it alongside the new one. What changed is that a second purpose became possible, and the lead purpose of the newest version moved with it.

The same paper sets out the design philosophy behind the presence purpose, and most of it still holds. Since the robot is a surrogate for human interaction, the default should be to replicate what a person would experience if they were there, starting with low latency and an image good enough to see details on the bench. It should have something equivalent to a head that can meet people face to face, ideally movable from desk height to standing head height, with head and base movements choreographed so that the operator's intent is visible. But it should capture the intent rather than copy the human literally: a screen with a face on it is better than a mechanical rubber face, which lands in the uncanny valley. And a robot need not stop at human abilities; it can see more, hear better or draw a straighter line than a hand can, to balance what it will never do, such as touch and smell. SEFI

Rule

Capture the intent, not the human. A face on a screen beats a mechanical face, and a robot is allowed to be better than a person at the things it can do.

V4's two screen images in chapter 02, asleep when nobody is connected and awake when someone is, are that rule in its simplest possible form.

02

Four robots before this one

Each generation removed something the previous one could not carry. The one that got used was the one that did the least.

It would be tidier to present V1 to V4 as steady progress. It was not. Each version was built by different people, usually as a thesis or a summer job, and each one inherited a machine that did not quite work and a pile of reasons why. What they have in common is more useful than what separates them: nearly every one of them made the robot simpler, and the working one is the simplest.

The generations at a glance

V0 is the student course prototype of 2019–2020 that preceded the robot the project later called V1. The book counts it as a precursor rather than a generation, which is why it speaks of five generations. They are generations of one idea, not refinements of one machine: several of them started almost from zero.

201920202021202220232024202520262027 V1 · plywood, swerve drive, arm V2 · rebuild of V1 V3 · simplified V4 · telepresence, in use V5 · brain works, body in design V0 project started
A precursor and five robots in seven years. The gaps are real: between 2021 and 2023, and again during 2024, the robot mostly waited for the next person to pick it up. Bars show active design and build periods, not service life. SEFIREQV2RHAKV4STAT
Five generations and their precursor: what actually changed
V0 · 2020V1 · 2021V2 · 2023V3 · 2025V4 · 2026V5 · in build
Purposetelepresence with an armtelepresence with an armsame, better finishjust move, remotelytelepresence, daily usetelepresence and assistant, mobile
Drivecar-like steering, rear differential3 modified swerve modules, 6 motorssame, stronger bearing housings2 DC motors + caster4 omni wheels, 4 motors4 omni wheels, brushless (BLDC) motors
ComputerIntel NUC + microcontrollerIntel NUC, ROS 2as V1ESP32 + Raspberry PiIntel NUC, plain PythonJetson AGX Orin 64 GB
Structuresquare aluminium profilebirch plywoodbirch plywood, gluedplywood floor, 3D-printed shellbirch plywoodin design
Mass—44 kg————
Statuscourse prototypeprototypeunfinishedprototypein useno body yet

References: V0 SEFI; V1 NUTJUV; V2 V2R; V3 V3; V4 V4; V5 STATRISK. Dashes mark figures that were never recorded; V2 and V3 have since been dismantled, so theirs cannot be recovered.

V0, 2019–2020: a course project with a sister robot

The first robot was not a thesis. It was built by student teams as the project work of four courses in mechanical and electrical engineering: analogue signal processing, electronics project work, systems engineering project work and technical design. The workshop staff acted as the client, setting the requirements from the user's point of view and leaving the students to re-examine them. The intention from the start was a relay: a modular, low-budget first version to prove the concept, then successive versions built by new groups on top of the previous ones' work. SEFI

The requirements that later became V1's list were first written here: a 90 cm door, aisles in a lab, a lift with a guardian, buttons and switches, two hours on a battery, recharging without special training, collision prevention, and a camera image good enough to see details of the work being guided. Opening doors was left out on purpose, as too demanding for a first robot.

Our case V0 · 2020 · how it was built SEFI

The frame was square aluminium profile, chosen because it is easy to machine, join and change when something goes wrong; the paper expected it to be replaced by “a less modular solution” later. The front wheels were steered like a car's by a stepper motor, with the steering angles set by the Ackermann condition (the geometry that turns the inner wheel more sharply than the outer), and the rear axle had a differential. The arm was an open-source BCN3D Moveo, about a metre long and meant to lift a kilogram, with its base at about 90 cm so it could reach table tops. The camera, a Ricoh Theta V on a pole at the top of the robot, saw 360 degrees. An Intel NUC, a small desktop computer, streamed it to a web page over WebRTC (the standard for real-time video in a browser), peer to peer, with the signalling server in the cloud, and passed drive commands over a serial line to a microcontroller. The web page let the operator rotate the 360-degree image with the mouse, drive with a joystick and move each arm axis with a pair of buttons.

Testing was planned for March 2020, and the pandemic delayed it. On 26 May 2020 the students instead drove the sister robot in San Francisco from Finland.

Block diagram of V0: a 360-degree camera, microphone and speakers on the main computer, which talks to the operator over a WebRTC peer-to-peer connection; a serial link to a motor-control microcontroller that drives the steering and drive motors, the manipulator, the batteries and the sensors.
V0 as the 2020 paper drew it. One main computer for video and audio, one microcontroller for everything that moves, a serial line between them, and the operator at the far end of a peer-to-peer video link. The split has survived in every version since; what changed was the number of motors on the right-hand side. No photograph of V0 has been found in the archive.Diagram: from the SEFI 2020 paper, by its authors SEFI

The paper's own conclusions are as useful now as then. Find out exactly where latency comes from before optimising anything, because it may have nothing to do with the robot's processors. Give students a stronger role as testers, not only builders, and collect user experience systematically between rounds. And the project raised a question that this book still has not settled: whether a workshop should teach rapid, iterative product development or manageable, clearly planned projects. A robot built by successive student teams cannot be run as a linear plan without skipping the development steps that make it useful. SEFI

Rule

In a student-built robot, make testing a role, not a phase. Someone has to use each version and write down what happened before the next team starts.

And answer the 2020 question openly at the start of every version: is this round rapid iteration, or a planned and bounded project? Either can work. Not deciding is what costs the most.

V1, 2020–2021: everything at once

V1 was designed as a master's thesis by Perttu Juvonen between September 2020 and March 2021, from the requirement list in chapter 01, continuing where the course project had left off. It is a remarkable piece of work for six months. The structure changed from aluminium to Nordic birch plywood, cut on a CNC router, chosen so that anyone with a router could build one and so that it could ship flat. The robot stood 137 cm tall on a 61 × 56 cm footprint, weighed 44 kg, and carried a Dorna 2 arm on a vertical carriage, an Intel RealSense depth camera and a 10-inch touch screen on a tilting neck, an Intel NUC running ROS 2 (Robot Operating System, a widely used software framework for robots), and a 24 Ah e-bike battery. NUTJUV

Render of the V1 design: a white plywood base on castors, a tall tower with a screen at the top and a dark robot arm on a carriage. Render of one V1 wheel module: a motor on top turns the module through a large gear to steer, and a second motor drives the wheel.
V1 as designed, and one of its three wheel modules. Each module steers and drives its own wheel, so three modules need six motors and a six-axis controller. The mechanism works; the software to coordinate it was never finished.Renders: V1 thesis work, 2020–2021 JUV

Omnidirectional movement was a requirement, and Juvonen chose a modified swerve drive: three wheel modules, each with one motor to steer and one to drive. That is six motors and a six-axis controller for a robot that mostly needs to go forward, turn and occasionally slide sideways through a door. His own conclusion in the thesis is the most useful sentence written about V1: that many features which are easy to implement mechanically need a significant amount of work on the software side, and that enough resources have to be allocated for it. JUV

Our case V1 · June 2021 · software status SWDEV

Three months after the mechanical design was finished, the software status note lists ROS 2 running under the Windows Subsystem for Linux, which could not see USB devices; an arm driver that existed only for ROS 1 and was being ported; sensor-board drivers that also existed only for ROS 1; and USB connection problems with the six-axis motor controller still being troubleshot. The mechanics were ahead of the software by about a year, and they stayed ahead.

Our case V1 · cost JUVREQ

The wish was under 5 000 dollars. The components came to just under 10 000 euros. The thesis gives the overrun as around 50 % in one paragraph and around 60 % in the next; both are recorded under Open items. The main single cause was the arm: the Dorna model originally specified became unavailable, and its replacement cost 3 500 dollars instead of 1 500. Most of the rest went to fast European distributors, chosen to keep a six-month thesis on schedule.

The V1 plywood base plate on the workshop floor during assembly, with a wheel module and loose wiring fitted.
V1 base during assembly, December 2020. Plywood, standard fasteners and parts that a student can replace with a screwdriver: that part of V1's philosophy has survived every generation since.Photograph: project archive

V2, summer 2023: making it look finished

V2 was a rebuild of V1 by two of JHC's summer research assistants, Jake Cumens and Jethro McLean, in the summer of 2023. Its brief was mostly visual: no visible screws, no laser soot, tidier joints, a centred screen, the speaker hidden. That brief was met. Rails were glued into recesses, the base cover got a new flex pattern, the charred edges were sanded away. The more interesting findings were the ones nobody asked for. V2R

Our case V2 · 2023 · a fatigue failure V2R

V1's wheel bearing housings were plywood, with many screws close together. The way the drive steered, dragging each wheel round to its new angle instead of letting it roll into the turn, put a repeated bending load through them, and the plywood fatigued. V2 replaced them with housings 3D-printed in carbon-fibre-filled nylon, shaped to stiffen the side plate, and recommended changing the steering so the wheel rolls while it turns. The nylon drive gears were also deforming.

The report then goes further than its brief. If there is a version 3, it says, it should have a new drive: three wheels with separate steering and drive motors complicate movement and need six motors where two would do. It lists the alternatives, tank drive with a caster, four-wheel tank drive, omnidirectional wheels, and says to choose by how much the robot actually needs to turn on the spot.

V2 also dropped the depth camera as unnecessary for what the robot was doing at the time, and noted that the arm's control unit probably had a faulty module. The rebuild was not fully finished. In parallel, a student, Robert Hämäläinen, wrote the control software for the V1/V2 drivetrain between 2023 and early 2024, talking to the motor controller over its native protocol and to a browser over WebSockets. CTRL

Rule

Choose the drivetrain by the motion the robot actually needs, and count the motors. Every actuator is a driver, a cable, a failure mode and a piece of software.

Omnidirectional motion is genuinely useful in a crowded workshop. It does not require six motors. Four omni or mecanum wheels (wheels with free rollers round the rim, so they can also roll sideways) on four fixed motors give the same freedom with no steering mechanism at all.

V3, 2024–2025: the simplest thing that moves

V3 came out of Sachintha Alwis Weerasinghe's 2024 master's thesis, which took the telepresence robot back to first principles with a systematic design method. It arrived at the simplest mobile Hynebot yet: an ESP32 microcontroller board driving two DC motors with Hall sensors (which count wheel revolutions) through a hobby H-bridge motor driver, a swivel caster at the front, remote control over MQTT (a lightweight messaging protocol), and video through an ordinary conferencing app on a tablet fixed to the robot. ALW

Miro Hakuli's 2025 thesis built a remote-control system for the V1-derived robot: a server between user and robot, socket.io (a library for live messaging between browser and server), and a browser interface. Its requirements came from interviews with the people around the project, and the one they stressed most was security: nobody should be able to move the robot without permission, and only one person at a time. Everything was tested except the motors themselves, because the robot had problems that prevented it. HAK

The V3 prototype from the 2024 thesis worked only on the local network. Real remote use, over a 4G connection from anywhere, was added at the end of 2025 as a JHC project, after Hakuli's thesis, together with a Raspberry Pi camera and live video and control in a web browser. With that, V3 did one thing, moving under remote control from somewhere else, and it did it. V3

The V3 base with its white lid removed, on a table at a seminar: a green 12-volt battery, two geared DC motors with white 3D-printed wheels, and a hand-soldered board carrying the microcontroller and motor driver, inside a plywood and 3D-printed shell marked HB v3.
V3 at the JHC Spring Seminar, 2024. Lid off: a 12 V battery, two geared DC motors, a hobby motor driver and the microcontroller on one hand-soldered board, 3D-printed wheels and shell on a plywood floor. Everything in the picture can be replaced by a student in an afternoon.Photograph: Teemu Leinonen, LUT University

V4, 2026: the one that got used

V4 is the version in service today. It kept the plywood tower and screen, and replaced almost everything underneath with the plainest parts that would do the job: a ready-made omnidirectional base the project received in 2024, with four omni wheels on four motors and two dual-channel Roboteq controllers; an Intel NUC running a handful of small Python services instead of ROS; three wide-angle USB cameras (forward, down at the floor in front of the robot, and rearwards); a USB microphone and speaker; a 24 V lithium battery; and a 4G router on board. It is driven from a web page, with keyboard control on a laptop and a touch joystick on a phone. The software was written in early 2026 largely by one staff member with an AI coding assistant, after the code left by several student generations proved too scattered to revive; chapter 10 says more about how AI assistance was used. V4

Our case V4 · the details that made it usable V4

A watchdog between the network and the wheels. The drive loop runs at 60 Hz, ramps every command over 0.15 s, and stops the motors if no command has arrived for 0.25 s. When the 4G link drops, the robot stops.

Access that answers the 2025 requirement. The controller page sits behind an authenticating tunnel; operators are admitted by e-mail address. The requirement from Hakuli's interviews, that nobody moves the robot without permission, is met by a service rather than by code we wrote.

Video sized for the link. 4G turned out to be the bottleneck, so the cameras stream their native compressed format at 640 × 480 and 15 frames per second. Object detection (YOLOv8 nano, 61 ms per frame on the NUC's CPU) runs in its own thread on a separate stream so that it can never delay the video used for driving.

A face that says whether anyone is there. When no operator is connected, the screen shows a sleeping face; when someone connects, it switches to the JHC logo. It is a two-image trick, and it answers the first question anyone near a telepresence robot has: is somebody watching me?

The idle screen: a sleeping emoji face with Z letters over the JHC logo. The active screen: the J. Hyneman Center logo, shown when an operator is connected.
V4's two faces. Left, nobody connected. Right, an operator is present. The switch follows the controller connection with about three seconds' delay. V4

Rule

A robot that is not used teaches nothing about what the robot should be. Get one version into daily use as early as possible, even if it does a fraction of what you want.

V4 does less than V1 was specified to do. It is the first Hynebot that people outside the build team have relied on, and most of what V5's requirements now say about latency, access control and status indication was learned from it.

What V5 inherits, and what it does not

In spring 2026 the decision was made to keep V4 exactly as it is, as a working telepresence robot, and to build V5 alongside it as a new platform that keeps the telepresence role and adds the assistant role. V5 inherits V4's drive kinematics and its drive, camera and display services as a starting point, the plywood-and-standard-parts philosophy from V1, the four-motor omni drive that V2's report pointed to, and the lesson from every version that the software is where the time goes. It does not inherit V1's arm-on-a-carriage layout, V1's middleware, or the assumption that the first body has to do everything. V4K

Part II

The brain: language on one edge computer

How V5 hears, thinks and answers without the cloud, and the measurement campaign that chose its language model. The studies had to be done before anything could be built, because nobody had measured how openly available models answer technical questions in Finnish on a computer small enough to ride in a robot. This part carries their complete results, including the ones that contradicted our first impressions.

03

First impressions do not survive measurement

One question asked once tells you about that question. The model we recommended after the first afternoon was not the one the data chose.

V5 had to speak Finnish and English. Roughly half of JHC's users are international students, and the other half deserve to be answered in their own language. The robot also had to run everything locally, for reasons chapter 07 sets out, which put the whole language model on one NVIDIA Jetson AGX Orin, an embedded computer built for robots, with 64 GB of memory shared between processor and graphics. So the question was concrete: which openly available language model, compressed to fit that memory, gives correct technical answers in both languages fast enough to talk to? V4K

None of that could be looked up. Published comparisons of language models are almost all in English, run on data-centre hardware, and score general knowledge rather than whether a model knows which way the filament goes into a particular printer. There was no table anywhere that said how a model compressed to fit a Jetson answers a safety question in Finnish. The choice also fixes everything downstream: the memory a model needs decides the computer, the computer decides the battery and the mass, and the time it takes to answer decides whether a conversation is possible at all. So the first months of V5 were measurement rather than building, and this part records what was measured and what it changed. Chapter 10 explains why the results are published here rather than in a journal first.

In plain terms

Open models, quantization and related terms

Open model
A language model whose weights can be downloaded and run on your own hardware, such as Llama, Gemma, Mistral, Qwen or Phi.
Ollama
A free program that downloads open models, runs them on your own computer and offers them to other software through a simple local interface. For V5 it is the piece that turns a downloaded model file into something the robot's own code can send a question to. It needs no cloud account. All our measurements went through it.
Parameters
The model's learned numbers. “70B” means about 70 billion. More parameters usually means more knowledge and more memory.
Quantization
Storing each parameter with fewer bits so the model fits in memory. Q4 means roughly four bits per parameter, Q3 roughly three. A 70B model is about 140 GB at 16 bits, about 42 GB at Q4 and about 32–34 GB at Q3. Every edge deployment makes this trade.
Mixture of experts (MoE)
A model built from many sub-networks of which only a few are used for each word it generates. It has the knowledge of a large model and the speed of a much smaller one. A conventional dense model uses all its parameters for every word.
Token
The unit a model reads and writes: a word or a piece of a word. Generation speed is given in tokens per second.
Context window
How many tokens the model can take into account at once. A larger window needs more memory.
Temperature
A setting for how much randomness the model uses when choosing each next word. At 0 it always takes the most likely word and gives the same answer every time; higher values let it take less likely words, which makes answers more varied and natural but less repeatable. 0.7 is a common default for conversation. We measured at 0.7 because that is how the robot will run. It is also why every question was asked three times, and why a repeat at temperature 0 is listed as open.

Our case V5 · 8 May 2026 · the first afternoon SEL

The first comparison was one question, “explain briefly how a 3D printer works”, asked once in each language to five models and scored by eye from 0 to 5. The English answers were all usable. The Finnish ones ranged from word salad (Mistral 7B invented words like “laatukone” and “lasimuotoilutekniikka”) to confidently wrong (Qwen 2.5 14B defined PLA as polystyrene). Mistral Small 24B gave the only good Finnish answer, and the note that day recommends it as V5's base model and hypothesises that Mistral is exceptionally good at Nordic languages.

The structured study four days later put Mistral Small 24B second of eight in Finnish accuracy, but clearly behind Gemma 2 27B (2.29 against 3.07) and with three hallucinations per Finnish answer. The early single-question tests had also marked Gemma 2 as unstable; with three repetitions it turned out to be the most repeatable model in Study A. Neither the recommendation nor the Nordic-language hypothesis survived.

Rule

Decide the protocol before you look at any answers: the questions, what a correct answer contains, how many repetitions, who scores and how. Then run it all.

The first afternoon was not wasted. It showed that the problem existed and roughly where. But the model choice for a robot that will answer safety questions has to rest on more than one prompt.

The protocol

We wrote 40 technical questions, each in Finnish and English, in six categories: general technical knowledge (10), JHC's own machines such as the Prusa MK4 printer, the laser cutter and the soldering station (10), scientific reasoning (5), workshop safety (5), everyday descriptive language (5) and demanding technical concepts such as G-code, mesh bed levelling and PID control (5). Before any model saw them, each question was given a list of facts an acceptable answer must contain and a list of red flags, claims that would be wrong. PLAN

Every model answered every question three times, at temperature 0.7, on the Jetson, through Ollama (both explained in the box above). The first run, Study A, covered eight models from 7B to 70B and took 28 hours 43 minutes over 12–13 May 2026: 1 920 answers, of which two timed out. A controlled quantization run added 480, and Study A2 at the end of May added 960 from a newer generation of models. In total 3 360 answers were generated on the robot's own computer. STAQ34A2

Questions
40 × 2
Repetitions
3
Models
11
Answers
3 360
Hardware
1 Jetson

Who scores 3 360 answers

Not a person. Every answer was scored blind by a separate, larger model acting as a judge: Claude Sonnet 4.6 through Anthropic's API, which saw the question, the answer, the acceptance criteria and the red flags, but not which model had answered. It scored factual accuracy, language fluency and completeness from 0 to 5 and listed every incorrect or invented claim it found. The prompt is reproduced in full below, because a judge is only as good as its instructions and anyone checking our numbers needs them. JUDGE

Evaluate the technical question-answer pair below. The answer was produced by a language model; assess it objectively. Output ONLY JSON, no other text, no markdown code blocks. QUESTION ({lang}): {question} ANSWER: {answer} ACCEPTANCE CRITERIA (what a good answer contains): {criteria} RED FLAGS (signs of hallucination): {red_flags} Score on a 0-5 scale: - factual_accuracy: factual correctness against the acceptance criteria (0=wrong, 5=accurate and complete) - language_fluency: fluency and grammaticality in the answer's language ({lang}) (0=incomprehensible/wrong language, 5=native-level) - answer_completeness: covers the essential points (0=does not answer, 5=complete) - hallucinations: list of incorrect/invented claims found (empty list if none) Write the reasoning field in English. Respond in exactly this JSON format: {"factual_accuracy": 0, "language_fluency": 0, "answer_completeness": 0, "hallucinations": [], "reasoning": "brief justification in English"}

Of the Study A answers, 1 918 were judged automatically, one was scored by hand after the judge declined a harmless question, and the two timeouts have no score. The failures we did see were mundane: the judge wrapping its JSON in a code block, or running out of output length on very long answers. Each was fixed and re-run, and every final score was checked to lie in the 0–5 range. STAT

Checking the judge, and then checking the humans

An automated judge has to be checked against people. We drew a stratified sample of 60 answers, ten at each accuracy level the judge had given and half in each language, and had three people score them for factual accuracy without seeing the judge's scores or the model names. One of them belonged to the project team; the other two were given minimal instructions and no hypotheses. IRR

In plain terms

Agreement between raters

Cohen's kappa (κ) measures how much two raters agree beyond what chance would give: 1 is perfect, 0 is chance. The plain version treats a 4-versus-5 disagreement as badly as 0-versus-5. The weighted version (κw, quadratic weights) counts near-misses as nearly right, which suits a 0–5 scale. Fleiss' kappa is the same idea for more than two raters at once.

Judge against three human raters, n = 60, factual accuracy
Pairweighted κcorrelation ragree on pass/fail (≥ 3)
judge · rater 10.570.6068 %
judge · rater 20.300.4062 %
judge · rater 30.730.7683 %
human · human, range0.51–0.680.57–0.74—
all four, Fleiss κ (unweighted)0.119

Source: IRRA1. Rater 3 had the most experience with the technical domain.

Read naively, a Fleiss κ of 0.12 is poor. Read properly, the judge agrees with the humans about as well as the humans agree with each other, and with the most experienced rater it reaches what is conventionally called substantial agreement. Published work on multilingual automated judges reports Fleiss κ around 0.3 on average and as low as 0.1 for some languages, so this is a hard task for everyone, not an unreliable judge. FULIU

Our case V5 · June–September 2026 · fluency bias IRRPS

The disagreements between the judge and the humans were not random. All three human raters were more lenient than the judge, by 0.3 to 1.3 points on average, and the biggest gaps were answers that read well and were wrong. In one question about linear and pressure advance in 3D printing, three different models described an industrial extruder that does not exist. The judge gave them zero, with reasons. The project team's own rater gave them three, because the text was fluent and sounded right. The same pattern was stronger in the most lenient rater.

We did not change our own scores afterwards to match the judge. Adjusting your ratings once you have seen the judge's destroys the check you are trying to make. We recruited more independent raters instead.

People are fooled by fluency, measurably

The failure that matters most for a robot in a workshop is not the answer that is obviously broken. It is the fluent, confident, specific answer that is wrong. Humans, including the person who wrote the questions, rate those answers too high. A demo that “sounds good” proves nothing about correctness. Chapter 06 comes back to this, because it is the strongest argument for how V5 answers machine questions.

04

Compression breaks the language before it breaks the model

One step of extra compression destroyed a large model's Finnish and left its English untouched. Guidance based on English benchmarks does not transfer.

Everyone who runs a large model on edge hardware quantizes it. The common guidance, that four bits per parameter costs very little quality, comes almost entirely from English benchmarks and from translation tasks at four bits. It rarely isolates one morphologically rich language, one that builds words from many endings as Finnish does, at three bits, which is where a 70B model has to go to fit comfortably in 64 GB. MARCHBORG

Study A contained exactly that case, and it was the most striking result in the data. The largest model, Llama 3.1 70B at three bits, wrote good English, judged at 3.48 for accuracy, and fell apart in Finnish at 0.48. Its Finnish answers were a quarter the length of its English ones, 18 % of them were under 100 characters or off-task, and at 1.8 tokens per second it took almost three minutes to produce each one. It was the worst model in the study in Finnish and the slowest. The collapse was deepest exactly where the robot needs it most: on the questions about JHC's own machines, its Finnish answers averaged 264 characters, the shortest of any model in any category. A1STA

That result alone cannot separate two explanations. Either this model family is weak in Finnish, or three-bit compression is the cause. So we ran the same 70B model again at Q3 and at Q4, on the same questions with the same settings, changing only the quantization.

012345 judged factual accuracy, 0–5 Finnish Q3 0.46 Q4 2.22 English Q3 3.56 · Q4 3.58 three-bit (Q3) four-bit (Q4)
Same model, same questions, same hardware. Raising precision from three to four bits almost quintupled Finnish accuracy and left English where it was. Llama 3.1 70B, 40 questions × 3 repetitions per language and level, context window 8 192 tokens in both runs. Q34A1
Llama 3.1 70B, Q3 against Q4, controlled run
FI accuracyFI fluencyEN accuracyEN fluencyFI broken output
Q30.461.743.564.6013 %
Q42.223.833.584.620 %

The effect is not carried by a few questions. Taking each question as the unit, so that three repetitions of one question are not counted as three pieces of evidence, 36 of the 40 Finnish questions improved from Q3 to Q4, one got worse and three were unchanged. The median improvement per question was 1.67 points (95 % bootstrap interval 1.00 to 2.33, a confidence range found by resampling the questions). The matched-pairs rank-biserial correlation, an effect size from −1 to 1 where 1 means every question improved, was 0.99 (0.97 to 1.00). In English the same comparison gave 8 better, 11 worse and 21 unchanged, a median change of zero and a correlation of −0.14 with an interval spanning zero: no effect. A1

Rule

Choose model size and quantization level together with the language the robot must speak, and test in that language. On constrained hardware a large, heavily compressed model can be worse than a smaller one in both quality and speed.

If your robot only ever speaks English you have far more compression headroom than we did. For us, Finnish is the constraint that sets the hardware envelope.

Two mechanisms, not one

The controlled run also separates two things that the first result mixed together. The collapse from 2.22 to 0.46 is caused by three-bit quantization, because nothing else changed. But even at four bits, the 70B model's Finnish (2.22) stays far below its English (3.58). That second gap is a property of the model family's Finnish, and quantization does not cause it and Q4 does not fix it. Other models in the study, several of them much smaller, did far better in Finnish at four bits (table in chapter 05).

Our case V5 · two Q3 runs that disagree A1Q34

The broad Study A run gave 18.3 % broken Finnish answers for the Q3 model; the controlled run gave 13.3 %. The only setting that differed was the context window: Study A used Ollama's default of 131 072 tokens, the controlled run 8 192, because the Q4 model will not load on the Jetson at the default. Its key-value cache, the memory the model keeps for the text so far, does not fit in memory alongside the 42 GB of weights, and it needs an explicit limit of 16 384 or less. Both Q3 figures are reported, and the Q3-against-Q4 comparison uses only the controlled pair.

Neither 70B variant is usable on the robot anyway. At Q4 an answer took about 113 seconds, at Q3 about 124.

What this does not show

This is one model family at one pair of quantization levels. The mechanism is consistent with published work showing that three bits is where measurable degradation begins and that the safe bit width depends on scale, but we have not shown that the threshold is general. CHANGKIM The two experiments that would strengthen it most are a smaller model at Q3 (Llama 3.1 8B at Q4 already scores only 0.61 in Finnish, which complicates the picture) and a repeat with greedy decoding (temperature 0), to separate the collapse from sampling noise at temperature 0.7. Both are listed under Open items.

05

Quality, speed and consistency are separate measurements

Our first selection criteria measured how much a model wrote, not whether it was right. Architecture mattered more than size.

The full table

Eleven models, two runs. Study A models answered in May 2026; Study A2 re-ran phi4 as a baseline alongside three newer models so that the two runs could be compared. A1A2

Blind-judged quality, 0–5, mean over 40 questions × 3 repetitions per language, ordered by Finnish accuracy
ModelRunFI acc.FI fluencyEN acc.EN fluencyhalluc. / answer
qwen3.6:27bA24.024.374.224.581.05
gemma3:27bA23.874.553.824.521.13
gemma4:26b (MoE)A23.824.374.244.500.86
gemma2:27bA3.074.173.444.871.62
mistral-small:24bA2.292.883.704.792.00
phi4 (14B)A22.242.913.674.702.58
phi4 (14B)A2.122.673.724.752.71
poro-34b-chatA1.773.831.914.263.53
qwen2.5:14bA1.211.523.504.733.35
llama3.1:8bA0.611.773.054.603.42
mistral:7bA0.580.893.074.673.70
llama3.1:70b Q3A0.481.723.484.672.42

Hallucinations are counted per answer with both languages pooled. All Study A models are four-bit except the 70B at three. References A1A2.

Three things stand out. English is easy: every model except the Finnish-specialised Poro scores between 3 and 4.3. Finnish separates the models by a factor of eight, and not in order of size. And the newest generation closed most of the gap: the three 2025–2026 models all score close to 4 in Finnish.

The criteria were wrong, and the data said so

Our case V5 · May–June 2026 · phi4 to gemma4 A2STAT

Before the judge had scored anything, phi4 looked like the winner of Study A. Its Finnish answers were as long as its English ones (length ratio 0.96), its length varied least from run to run after Gemma 2, and it produced 15 tokens per second. The protocol's decision rules for replacing it had been written around those numbers: a new model must answer in under 15 seconds, match phi4's Finnish-to-English length ratio and match its consistency.

The judge then showed that phi4 writes Finnish of the same length as its English and of much lower quality: accuracy 2.24, and 2.58 hallucinations per answer. Applied mechanically, the rules kept phi4 against every challenger, because the latency limit was unreachable on this hardware for every model including phi4 itself, and because a length ratio measures the balance of text volume, not of language quality. We chose gemma4:26b on what the data meant, not on the rules as written: roughly 1.7 times phi4's Finnish accuracy, a third of its hallucinations, faster and more consistent, with no English mixed into any of its Finnish answers.

Rule

Length, speed and consistency are not quality. Measure correctness directly, and write decision criteria that a model could actually pass on your hardware.

When the criteria and the data disagree, say so, change the criteria openly, and record why. Silently applying a rule you know is wrong is worse than breaking it in the open.

Architecture beat size

For a robot, the answer has to arrive while the person is still standing there. In the same protocol the two newer dense 27B models took around two minutes of generation per answer on the Jetson: gemma3:27b averaged 126 seconds and qwen3.6:27b 138. The mixture-of-experts gemma4:26b, nominally the same size and with the best overall quality, averaged 29 seconds on the same long answers. Picking the biggest model that fits in memory is the wrong instinct on edge hardware; pick the architecture that uses only part of itself for each word. REPA2

Our case V5 · 9 June 2026 · warm and cold RAG

Those averages looked alarming for a conversational robot until we measured the case that matters. The study timed long answers, and the first answer after a pause includes 15–20 seconds of loading the model's weights into memory. With gemma4 kept loaded and asked short workshop questions, the model answered in 2.2–3.1 seconds, 2.5 on average, and in 3.4–4.7 seconds when it first had to look up the manual. The operational consequence is simple: the model is loaded at start-up and kept resident. The 2.5 seconds is the model alone. The whole path from the end of speech to the first spoken word has not yet been measured as one number, and chapter 07 lists it as open.

Consistency is its own axis

A robot that gives a different answer each time you ask erodes trust even if every answer is acceptable. We measured it two ways across the three repetitions: how much the answer length varies (the coefficient of variation, standard deviation as a percentage of the mean) and how much the wording overlaps (ROUGE-L, the share of words that appear in the same order in two answers, from 0 to 1). The two agree well (rank correlation, Spearman ρ, of −0.79 across the eight Study A models), and neither follows model size. REP

Run-to-run variation in answer length, coefficient of variation in % (lower is more consistent); ROUGE-L 0–1 (higher is more consistent)
ModelallFinnishEnglishROUGE-L
gemma4:26b5.75.95.40.38
gemma3:27b7.78.17.30.34
gemma2:27b11.512.210.80.30
phi412.412.512.30.25
qwen3.6:27b13.012.913.10.29
qwen2.5:14b14.916.313.40.22
mistral-small:24b15.919.312.50.27
poro-34b-chat21.019.622.40.24
mistral:7b30.843.118.40.22
llama3.1:8b34.052.615.40.20
llama3.1:70b Q338.464.312.50.21

The instability sits almost entirely in Finnish. In English every model varies by 5–22 %; in Finnish the range is 6–64 %. An evaluation run only in English would not have seen the problem at all. The two measures can also disagree in a way that matters. phi4 is the second most consistent Study A model by answer length, yet its judged accuracy varies more from one repetition to the next than any other Study A model's (standard deviation 0.49 points against 0.29–0.48 for the rest): it says the same amount each time, but not the same thing. And the controlled 70B run shows that quantization explains part of the instability but not all of it. Going from Q3 to Q4 reduced the length variation from 41.9 % to 30.3 %, still far above the stable models. REP

The phi4 baseline, run twice three weeks apart after an update of the model runtime, reproduced to within 0.3 points on every measure, which is the evidence that the two runs can be compared at all. A2

What changed between generations

Two families appear in both runs. Qwen went from 1.21 in Finnish accuracy (2.5, 14B) to 4.02 (3.6, 27B). That jump is real but it mixes a new generation with a model twice the size, so it is not a clean generational effect. Gemma, at a constant 26–27B, went from 3.07 to 3.87 to 3.82 across three generations, with hallucinations falling from 1.62 to 1.13 to 0.86 per answer. The practical reading is that a Finnish-speaking robot built in 2024 would have needed a Finnish-specialised model or a larger computer, and one built in 2026 does not. A1A2

Our case V5 · what we set aside, and why STAT

Poro 2, a Finnish-trained model family, was planned for A2 and dropped on 28 May: getting it onto the Jetson required a format conversion and a custom build of the inference engine, and by then general models were already better in Finnish than the Poro 34B chat model we had tested. qwen3:30b-a3b was removed from A2 because it kept writing English reasoning into its answers even with reasoning switched off, so its timings were not comparable. A second identical robot running a different model each month, proposed in an early plan, was dropped because it would have confounded model, time and usage, and the offline studies answered the question with better control. None of these is ruled out for ever. Each was set aside for a stated reason that can be revisited.

06

Retrieval is a safety feature

A robot standing next to a machine and answering fluently and wrongly about that machine is a hazard. Letting it read the manual first fixes most of that.

In plain terms

Retrieval-augmented generation (RAG)

Before answering, the system searches a set of documents, here the workshop's machine manuals, for the passages most relevant to the question, and gives them to the language model together with the question. The model answers from the passages rather than from memory. The search works on meaning rather than exact words, using an embedding model that turns text into vectors, so a Finnish question can find an English passage.

V5's retrieval runs on the Jetson: the multilingual-e5-base embedding model, a local ChromaDB vector store (a database searched by embedding), and as the first knowledge base the Prusa MK4 printer handbook, 67 pages indexed one page per chunk. A Finnish question and its English equivalent land very close together in the embedding space (cosine similarity 0.939, where 1 would mean the same direction), so Finnish questions reliably find the right page in an English manual. Across six probe queries, three topics in two languages, the correct or topically correct page was in the top two results every time. That matters, because almost all machine documentation is in English and half the questions will not be. RAGIA2D

The first retrieval prototype, built with Mistral Small 24B before gemma4 was chosen, showed the qualitative pattern that the later test confirmed. With the manual, answers contained procedures that exist only in that manual: cleaning the print sheet with isopropyl alcohol, running mesh bed levelling, confirming with the rotary knob. Asked something the manual does not cover, the grounded model declined instead of inventing an answer. Without the manual it did neither. A2D

Our case V5 · 9 June 2026 · with and without the manual RAG

Six questions about the MK4, each answered by gemma4 with and without the manual. With the manual it won four, lost one and drew one. Two of the wins are the reason retrieval is on by default for every machine question. Without the manual, the model stated the printer's build volume as 200 mm cubed; the handbook says 250 × 210 × 220 mm. And asked about the MK4's load cell, it explained confidently that it measured material strength. The load cell measures the distance between nozzle and bed. Both answers were fluent, specific and plausible. Both were corrected the moment the model could read the page.

The loss was instructive too. The retrieved page said, in effect, “see the online help”, and the model passed that on instead of answering. A retrieved passage that points elsewhere needs handling in the prompt.

This is the same failure the human raters showed in chapter 03, one layer up. There, people scored fluent nonsense too high. Here, the model produces fluent nonsense about the very machine the student is standing at. An earlier comparison in the project had already shown that a 24B model with the manual beats a 70B model without it on device-specific questions. RAGI That finding is about quality. The load-cell answer is about safety, and it is the one that decided the design.

Retrieval substitutes for scale

Put the two findings of this part side by side and they point the same way. A large model compressed to fit the edge computer loses the minority language first (chapter 04). A smaller model that can read the documentation answers machine-specific questions that neither the large model nor its own memory can. On a fixed memory budget, capacity spent on retrieval buys more than capacity spent on a bigger model. That is the practical form, for a bilingual workshop, of the broader argument that small, suitably augmented language models are enough for most agent-like tasks. BELCAK

What retrieval does not fix

The prototype also showed three things retrieval cannot correct, and they are as important as what it can. A2D

  • Antonyms sit next to each other. A Finnish question about loading filament retrieved the unloading page and the loading page at almost identical distances, 0.415 and 0.416, because the two pages share nearly all their context. Meaning-based search is weakest exactly where the meaning is opposite and the words are the same.
  • Translation happens at answer time. The manual is English and the answer is Finnish, so the model translates as it answers, and it rendered at least one control name as a non-standard Finnish compound that appears nowhere in the source. Grounding fixes facts, not terminology.
  • It trusts what it is given. If speech recognition mishears the question, retrieval faithfully answers the wrong one. Chapter 07 follows that failure through the whole chain.

Rule

Route every question about a specific machine through that machine's documentation, show where the answer came from, and send anything about hazardous operations to a person.

Retrieval makes up for model size, as shown above. More importantly, it turns an unverifiable answer into one that cites a page.

Six questions is a smoke test, not a result

The comparison above is twelve answers. It is enough to decide that retrieval stays on; it is not enough to publish a retrieval accuracy. A proper evaluation needs at least twenty annotated questions with the correct manual page marked for each, three repetitions, both languages, and the judge. The knowledge base also needs to grow beyond one printer to the laser cutter, the soldering station and the measuring instruments. Both are open.

07

Most of the work is the pipeline around the model

Speech in, speech out, nothing sent anywhere, and every stage in between can fail on its own or pass a mistake on to the next.

one NVIDIA Jetson AGX Orin 64 GB · no network needed MicUSB Hearingfaster-whisper Routermanual or not Manualse5 + ChromaDB Modelgemma4:26b Voicepiper Speaker+ screen model: 2.5 s warm 3.9 s with manual end of speech → first spoken word: not yet measured · target under 5 s
V5's brain, as it runs on the desk today. Only the language-model step has a measured latency. The target for the whole path is the first spoken word within five seconds of the user finishing. STATRAG

Why nothing goes to the cloud

Three reasons, in order of weight. Privacy: a microphone in an open workshop hears people who never chose to talk to a robot. If audio never leaves the device and is never stored, most of the privacy problem does not arise (chapter 09). Independence: the university network blocks some services, the robot moves between networks, and a robot that stops talking when the connection drops is not a robot anyone relies on. The point of the exercise: showing that a capable, bilingual assistant fits on one board is itself part of what the workshop is demonstrating. The one exception during development was the automated judge in chapter 03, which ran on stored text answers, never on audio, and is not part of the robot.

Hearing and speaking

Speech recognition is faster-whisper with the small model on the Jetson's GPU. It does not install from the standard package on the Jetson: its engine, CTranslate2, had to be compiled from source with CUDA support (NVIDIA's platform for computing on the GPU). Speech synthesis is piper, with a male Finnish voice and a female English one. That asymmetry is not a design choice. There is no good open female Finnish voice, which constrains how the robot presents itself. STATPS

Our case V5 · a word the robot could not hear STAT

The small Whisper model turns “3D-tulostin”, the Finnish for 3D printer and the single most common noun in our workshop, into something like “kolmendeetuloistin”. A recognition error on the key noun reaches retrieval as a different word, and a confident answer to a slightly different question is worse than no answer. Testing the medium model for its Finnish word error rate, the share of words recognised wrongly, is on the list; so is a proper wake word, because listening for five seconds at a time does not work in a noisy workshop.

Errors travel down the chain

Each stage of the pipeline has its own benchmark: word error rate for recognition, precision for retrieval, judged accuracy for the model. Each benchmark assumes the stage receives clean input. In a working robot none of them does, and the failures we saw on the desk were not in any one stage. They appeared where the stages meet. A3D

Our case V5 · desk prototype · three cross-stage failures A3D

Error laundering. When recognition turned “3D-tulostin” into a nonsense word, the model did not flag it. It built the nonsense word into a fluent, confident answer grounded in the manual. The grounding meant to make answers trustworthy turned a transcription error into an authoritative statement. The recognition benchmark counts that error once, in its own stage, and says nothing about what it cost downstream.

Grounding masks instead of correcting. Retrieval is usually justified as protection against invented answers. It protects against the model's own inventions, not against mistakes made before it: retrieval and model both work faithfully on a corrupted transcript, and because the answer cites the manual it looks more reliable, not less.

The gap between source and answer language. The translation error in chapter 06 exists in neither the English source nor any English-side measurement. No single-stage metric can see it.

The mitigations follow from the failures and each can be tested: recognition that reports its confidence and flags uncertain domain terms before they reach the model; a prompt that makes the model say when a word in the question looks wrong; re-ranking or hybrid word-and-meaning search to separate near-antonyms; and translating the retrieved passage as a separate step before the answer is formed. A3D

Rule

Log every intermediate result at every stage boundary, for every question: the transcript, the retrieved pages and their distances, the answer, the timings. Only then can a wrong answer be traced to the stage that caused it.

The measurements that would turn these observations into rates are defined and not yet made: word error rate by language and model size, retrieval precision and rank, how often a domain-term error reaches the final answer, what share of wrong answers each stage causes, and a latency budget per stage. They need the robot in use, and the interaction log is designed to collect them as it is used.

Traps worth knowing about

A switch that only hides what it switches off

Several recent models “think” before answering, generating a long internal monologue. For a robot that is pure latency. In the Ollama version we used, the option that looks like it disables thinking inside the model options only moves the monologue into a separate field: the model still generates it, and the token count stays in the thousands. Only the top-level request field "think": false actually stops it. We caught this with a five-minute test before a forty-hour run; it would otherwise have silently corrupted the whole study.

Rule

Before any long run, run a five-minute version of it and read the raw output, not the summary.

The Jetson is not a desktop

Never install desktop NVIDIA driver packages, server variants or the generic cuDNN package on a Jetson: they break the CUDA installation, and an unattended apt autoremove afterwards will remove the toolkit entirely. That cost a day and a reinstall of 1.8 GB of packages. The depth-camera library and PyTorch-adjacent packages also have to be built from source or taken from NVIDIA's Jetson-specific builds. Lock the environment, record every version, and keep a tested restore procedure.

The rest of the software is deliberately plain. A router module decides per question whether to search the manuals, keeps the model loaded, and sets the thinking switch correctly for each model. Development happens on a laptop, code goes through a GitLab repository to the Jetson, and the robot's own network is kept separate from the university's production network, with the development machine between the two. ROUT

Part III

The body, and what has to come before it

V5's platform is largely still being designed. This part records what is decided, what is not, and the safety and privacy work that sets requirements for the frame before it is drawn.

08

The drivetrain is chosen; its controllers are not yet proven

An omnidirectional base built from parts already on the shelf, sensing that sees all the way round, and an arm for gestures. The body above the base is not described here.

V5's body is still being designed, and this record does not describe its structure above the base. What is settled is the part JHC has worked out on its own account: the drivetrain, the sensing the base needs, and what the arm is for.

The drivetrain

V5 uses parts already on the shelf: four Nanotec DB43 brushless motors with 4:1 planetary gearboxes, 100 mm omni wheels, and four Electromen EM-356A-SBL motor controllers driven from the Jetson over a single RS-485 serial bus using Modbus, a simple industrial communication protocol. Raw geometry gives about 3.9 m/s at full motor speed, far too fast indoors, so the speed is limited in the controllers themselves, below the reach of any higher software layer. V4's omni-wheel kinematics carry over unchanged; only the motor-controller driver has to be rewritten. PSOV

Our case V5 · August 2026 · the controllers are positioning drives EMX

Reading the controller's register map while writing the driver showed that the EM-356A-SBL is designed to move a linear actuator to a position between end stops. It has no speed set-point on the bus. An omni base needs four wheel speeds, not four positions. Speed control can be synthesised by setting a far-away target position and adjusting the controller's maximum-speed parameters in RAM, but that is a workaround, with a position counter that will eventually wrap and a servo loop tuned for positioning rather than smooth driving. The decision rule is a bench test: one motor, one controller, speed control from the Jetson, before any frame is drawn around them. If it does not behave, a speed-controlled brushless driver replaces the controllers, and the motors and gearboxes stay.

Rule

Before you design a frame around a controller, make one motor turn from your own computer, the way the robot will use it. Datasheets describe what a controller was designed for, not whether it does what you need.

Sensing, and the order the frame has to respect

An omni base moves sideways and backwards as readily as forwards, so the distance sensing has to see all the way round. A 270-degree scanner leaves a blind sector exactly where the robot can drive. V5 will carry a 360-degree navigation LiDAR, a laser scanner that measures distances all round, from the start, with a safety-rated scanner as a later addition, not a replacement. The frame is being drawn so that adding one later costs a bracket, not a rebuild: an Ethernet and 24 V route to the scanner mount, a swappable mounting plate, a reserved mass and volume, and a free dual-channel safety relay input. OVSTAT

Rule

Some decisions are cheap before the frame is drawn and expensive after. Settle them first: battery and computer low for stability; a clear 360-degree line of sight for the scanner at obstacle height; the depth camera looking forward and slightly down; emergency stop, main switch and charging point designed in, not added.

The arm: gestures, not manipulation

V1 was specified around an arm, and V5 still has one: the Dorna Arm 2. Its software is ready. The key discovery was that its home position, with every joint resting against its hard stop, is not all zeros but 180, 180, −142, 135 and 0 degrees, which explained every alarm in the first tests. A safe parking pose one degree off the stops was established so that nothing drops when power is cut, and a complete demonstration sequence has been validated. DORNA

In V5's first year the arm makes gestures and nothing more. Handing people objects needs a gripper, and an industrial electric gripper weighs around a kilogram, a large share of what a small desktop arm can carry. Carrying things on a tray or a rack does the same useful job with a fraction of the risk. STAT

09

Safety and privacy are design inputs, not paperwork

The risk assessment decides the emergency-stop topology, which is an input to the electrical design. The privacy memo decides that audio is never stored, which is an input to the software.

Safety was on the first requirement list in 2020. The course project already required the robot to prevent collisions on its own, with pedestrians and glass doors named explicitly, and its paper suggested a pleasant sound whenever the robot moves, so that people nearby know even if they are not looking, and geofencing, a virtual boundary, to keep remote operators out of places the robot should not go. SEFI The sound is now a line in V5's risk assessment. V1 added a table of hazards and mitigations mapped to ISO 12100, ISO 10218, ISO 13482 and ISO/TS 15066, and a rule that the robot stays under 60 V DC. No single standard covers a hand-built telepresence robot with an arm; Juvonen's thesis applied the relevant parts of several. That approach has not changed. The 2020 paper also asked a question that belongs next to the hazards: can a telepresence robot be misused in an educational setting, and how would that be prevented? SAFE1JUV

For V5, the risk assessment was written in June 2026, before the platform was designed, so that it could set requirements instead of auditing a finished machine. It identifies ten hazards across three phases (the desk unit, the mobile base, and the arm on the base) and rates each by likelihood and severity before and after mitigation. RISK

V5 risk assessment, the hazards that shape the design
HazardRisk beforeMain measuresRisk after
Collision with a person4speed limit in the motor controllers; dead-man remote control; motors stop if the control link is lost; padded edges; sound on moving off2
Tipping over6low centre of gravity; 10° tilt test before use; floor survey; no ramps3
Battery fire3LiFePO4 (lithium iron phosphate) chemistry; battery management system (BMS) with voltage, current and temperature protection; main fuse and switch; charging only at a marked, supervised place3
Arm pinch or impact4only pre-programmed, validated gestures; speed and range limited in the controller; stops on e-stop2
Wrong technical advice leads to an accident at a machine4machine questions always answered from the manual, with the source page; hazardous operations always referred to staff; the model is instructed never to suggest bypassing guards2

Scores are likelihood × severity, each 1–3. The full table has ten hazards. The battery's residual risk is accepted because its severity cannot be reduced technically; it is managed by prevention. RISK

The last row is the one a conventional machine-safety assessment would not contain, and it is why chapter 06 exists. A robot that gives advice at a machine can cause an accident without moving at all.

The emergency stop is wired

The emergency stop cuts motor power in hardware: a red, latching button on the robot, and a normally-closed loop through every moving or hot accessory. It is never wireless, never over the network and never in software. A software stop exists as well, never instead. For a machine moving among students in the EU this is not negotiable, and it applies to every accessory that is ever bolted onto the robot. RISK

Privacy: the microphone hears everyone

A robot in an open workshop hears people who never chose to speak to it. V5's privacy memo turns that into system requirements rather than a policy. GDPR

  • Audio is never stored. Speech is transcribed on the device and the audio buffer is discarded. This is a system requirement, not a habit.
  • No identification. No speaker recognition and no face recognition; the camera detects object classes such as “person” in memory and keeps no frames.
  • Logs describe interactions, not people. The interaction log records the question text, language, route, latencies and the manual page used, with no user identifier, and it stays on the device or in university-managed backup, never in the code repository.
  • It listens only when asked. A wake word starts recognition, and the robot shows whether it is listening. Signs in Finnish and English say what is processed and where to read more.
  • Research is separate. Any study with students uses informed consent, an ethics pre-check and its own data management plan.

Retention and legal basis are proposals awaiting the university's data protection officer, and the risk assessment awaits the occupational safety organisation. Both are listed as open.

Rule

Write the risk assessment and the privacy memo before the frame and the software, while their conclusions are still cheap to follow.

Part IV

Method, and the road ahead

How the work was done, why it is published in this form, what that costs in rigour, and what comes next.

10

The method is part of the result

AI assistants were used throughout. The rules that kept the numbers honest were the same ones we would want from a human colleague.

AI assistance, openly

Much of V5's software, the analysis scripts, the study write-ups and this record were written with AI assistants: several general-purpose language models, from more than one provider, used as coding, analysis and writing tools alongside the project team. The automated judge in chapter 03 is a separate matter: one named model, used only to score answers, and documented there. AI use is stated plainly because it bears on how the work should be read. An assistant is fast, tireless and often right; it is also fluent when it is wrong, which is exactly the failure this book warns about. The working rules that follow exist because of that.

  • Every citation and every gap claim is checked against the original paper, not taken from an assistant's summary. Some were wrong on first draft: a citation pointing to the wrong authors, and two different papers merged into one.
  • Every number in the result tables is recomputed from the raw data by one script, and a second script compares every table cell in the drafts with the recomputed value before anything is published.
  • Human ratings are never changed after seeing the judge. Disagreement is resolved by adding raters, not by adjusting the ones you have.
  • Null and single-model results are not stretched. The quantization threshold in chapter 04 is claimed for one model family, because that is what was tested.
  • Five-minute smoke tests before long runs, and the raw output read by a person.

The open use of AI is also part of what the project is meant to teach. Students will work with these tools whatever the curriculum says, and the useful lesson is not whether to use them but how. Hynebot is meant to show it in practice: a language model working side by side with the people on the project, writing code, running analyses and drafting text far faster than they could alone, while the decisions stay with people. The model proposes; a person chooses the robot's purpose, sets the protocol, reads the raw output and signs off every number. The same holds for what is learned. A model can produce working code for a motor controller, but understanding why the controller turned out to be a positioning drive, or why a fluent answer can be wrong, is what a student takes away from the project, and that cannot be handed over.

Rule

Let the AI do the work it is fast at, and keep the decisions, the checking and the understanding with people.

If a student cannot explain a piece of the robot that an AI wrote, that piece is not finished.

Why the results are published here

The measurement campaign was written up as four studies, and their content is in this book:

The four studies and where they are here
StudyChaptersData
When quantization breaks a minority language03, 04, 05complete
Reproducibility does not track scale05complete
Retrieval over scale06qualitative quantitative evaluation needs data from use
Where the pipeline breaks06, 07qualitative propagation rates need data from use

Two of the four have complete data; the other two are incomplete. Retrieval over scale has a mechanism, a cross-language probe and a smoke test, but no measured retrieval precision. Where the pipeline breaks has three documented failure modes, but no rates: how often a misheard word reaches the final answer, or what share of wrong answers each stage causes. Both are incomplete for the same reason: the data they need comes from the robot in use, through its interaction log, and the robot is not yet in use.

We have chosen to publish the results here, openly and with their sources, rather than take them through journal peer review first. That is not a judgement on peer review, which this work would benefit from and may still get. It is a question of what the studies were for. They were not research for its own sake. They were the groundwork a physical robot needed before it could be built: which model, at what compression, with what retrieval, on what hardware. Those answers were needed within weeks, by the people building V5, and they are needed this year by the next team, by other workshops building robots with students, and by the companies around JHC deciding what to build. A public, versioned record lets the work be corrected and iterated as freely and as quickly as the robot itself.

For groundwork of this kind the review that matters most is the machine. If the robot built on these findings answers correctly, in Finnish, at the printer, within the time a person is willing to wait, then the theory and the plans were right. If it does not, the record shows exactly which number to go back to. The cost is that nobody outside the project has yet checked these results formally. We have tried to compensate by publishing what a reviewer would ask for: the question design, the judge prompt verbatim, the rater agreement with its weaknesses, question-level statistics rather than inflated sample sizes, the runs that went wrong, and the limits below. The questions, raw answers, scores and analysis scripts are available from the J. Hyneman Center on request.

Limits of the language-model results

  • One set of 40 questions, written by one domain expert and not piloted.
  • One automated judge, from a single model family, checked against three human raters on 60 answers.
  • Sampling at temperature 0.7 only; no greedy-decoding control.
  • One hardware platform and one runtime; latencies are specific to both.
  • The quantization threshold shown for one 70B model at Q3 and Q4.
  • The retrieval comparison is twelve answers about one printer.
  • Measured parts of a system, not a system in use.
11

Each step must leave something usable

The biggest unknown is not technical. Nobody who did not build V5 has used it yet.

The plan is a set of layers, each of which has to be usable on its own before the next is added. The brain and the body are separate layers, which is what lets the brain be finished on a desk while the body waits for parts. STATPS

V5 in layers
LayerWhat it addsStatus, October 2026
1 · Deskthe speech, retrieval and language pipeline in a fixed station beside the 3D printers; wake word, services that restart themselves, interaction logbrain works station not yet deployed
2 · Partsscanner, battery, frame material, bought on JHC's own fundingin procurement
3 · Mobilethe base under remote control, V4's drive software ported to the Jetson, the assistant on top; drives and carriesin design
4 · Comes when calledturns to and drives towards a marker or a person, stops for obstacles; the arm gestures. No map.target end of 2026
5 · Free navigationmapping, localisation, route planning anywhere in the buildingbehind a decision gate

Layer 5 opens only when layers 1 to 4 work and the interaction logs show that people actually need the robot to go somewhere on its own. Until then it stays closed, deliberately.

Rule

Plan so that each step leaves something people can use, even if the next step never happens.

A

Checklist: starting a workshop robot

The questions we would answer first if we started again. None of them names a part.

  • Which purpose leads: presence, assistance or demonstration? Written down, with the test that tells you it works.
  • What is the smallest version that someone outside the team would use every week? Build that first.
  • How many actuators does the motion you need actually require? Count the motors before choosing the drive.
  • Can every structural part be bought locally and repaired by the next student?
  • Does one motor turn from your own computer, the way the robot will drive it, before the frame is drawn?
  • What stops the wheels when the network drops, and how fast?
  • Is the emergency stop wired in hardware, through everything that moves or heats?
  • Where do the battery, the computer and the distance sensor go, before anything else is placed?
  • Which languages must it speak? Test every model and compression level in those languages, not in English.
  • Is the evaluation protocol fixed before you look at any answers, with repetitions and a scorer you have checked against people?
  • Do answers about a specific machine come from that machine's documentation, with the source shown?
  • What does the microphone keep, and who can see it?
  • Is autonomy behind a gate that real use has to open?
  • Which subsystem got the hours last month, and is that the one that leads?
B

Open items and contested figures

Where the project's own sources disagree, or a decision is still open. Listed rather than resolved silently.

Open items (O) and editorial gaps (E)
#ItemWhat the sources sayStatus
O1V5 indoor speed limitRisk assessment: 0.5 m/s, set in the controllers RISK. Progress summary and drivetrain notes: about 1 m/s PS.decide the risk assessment governs until changed
O2V5 motor controllersEM-356A-SBL chosen 20 Aug 2026; found to be positioning drives 25 Aug. A speed-controlled brushless driver is the fallback. EMXbench test
O3V5 navigation scannerEarly plan: RPLIDAR A2M12 V4K; later: RPLIDAR S2, with a safety scanner as a future addition OV. Requirement: full 360° coverage.procure
O4V1 budget overrunThe same thesis gives “around 50 %” and “around 60 %” JUV.recorded, not resolved
O5Hallucination countsSome internal notes label the pooled per-answer counts (gemma4 0.86, phi4 2.58) as Finnish-only. The computed tables pool both languages; this book uses them as pooled. A2ROUTcorrected here
O6gemma4 generation time29 s is described as a median in some notes and computed as a mean in the analysis. This book says average. REPto verify from raw data
O7Q3 broken-output rate18.3 % (Study A, context 131 072) and 13.3 % (controlled run, context 8 192). Both reported; explained in chapter 04. A1reconciled
O8Quantization threshold generalityShown for Llama 3.1 70B only. Needed: a smaller model at Q3, and a greedy-decoding control. A1experiment
O9End-to-end latencyOnly the language-model step is measured (2.5–3.9 s warm). Target for first spoken word: under 5 s. STATmeasure
O10Retrieval evaluationTwelve answers, one printer manual. Needed: ≥ 20 annotated questions, both languages, more machines. RAGmeasure in use
O11ApprovalsRisk assessment awaits LUT occupational safety; privacy memo awaits the data protection officer. RISKGDPRpending
O12Development laptop memoryThe early V5 plan specifies 64 GB V4K; the machine bought has 48 GB STAT.recorded
E1History, V2–V3Exact dates of the V2 and V3 work are reconstructed from documents. V3 originated in the 2024 thesis ALW; 4G remote use was added at the end of 2025 as a JHC project, after the 2025 thesis HAK (confirmed by the author, 5 Oct 2026).editorial
E2V0 detailsTaken from the 2020 conference paper. Its underlying source, the course final report by Rahikainen, Tuimala, Putkonen, Salminen and Grasbeck (2020, in Finnish), has not yet been located; no photograph of V0 has been found, and its mass is unknown.editorial
R

References

Every reference key in the text points to an entry in this list, grouped by kind of evidence. Paths refer to the project archive: “V1–V4” is the folder 06 Hynebot, “V5” is 07 Hynebot V5.

Our own measurements

STA

Study A: eight-model comparison, raw data

V5/02 Tutkimus/Raakadata/study_a_results.json · Analyysit/study_a_tulokset.md

1 920 answers, 12–13 May 2026, 28 h 43 min on the Jetson AGX Orin. Protocol, model list, answer lengths and speeds.

Q34

Controlled Q3 against Q4 run

V5/02 Tutkimus/Raakadata/study_a_q3q4_ctx8192.json · Arvioidut/study_a_q3q4_judged.json

480 answers from Llama 3.1 70B at Q3 and Q4, context 8 192 tokens in both; 479 judged.

A1

Computed results for the quantization study

V5/02 Tutkimus/Analyysit/artikkeli1_luvut.md · q3q4_tilastot.json · Artikkeli 1/artikkeli1_v7.tex

All quality figures for the eight Study A models, question-level statistics with bootstrap intervals, and the rater-agreement figures, recomputed from raw data by one script.

A2

Study A2: generation comparison

V5/02 Tutkimus/Analyysit/generation_comparison.md

960 answers, end of May 2026: phi4 baseline, gemma3:27b, gemma4:26b, qwen3.6:27b. Judged with the same protocol.

REP

Computed results for the consistency study

V5/02 Tutkimus/Analyysit/artikkeli4_luvut.md · Artikkeli 4/artikkeli4_v3.tex

Length variation and ROUGE-L across repetitions for all models; mean generation times for A2.

JUDGE

Judge script and prompt

V5/02 Tutkimus/Skriptit/judge_study_a.py

The verbatim prompt reproduced in chapter 03. Judge model: Claude Sonnet 4.6.

IRR

Inter-rater files

V5/02 Tutkimus/IRR-arviointi/

60-answer stratified sample, answer key and three independent human ratings.

RAG

Retrieval comparison, 9 June 2026

V5 project status, section 12 · rag_vertailu_tulos.json (on the Jetson)

Six Prusa MK4 questions with and without retrieval; warm latencies.

RAGI

Retrieval implementation notes

V5/02 Tutkimus/Analyysit/rag_toteutus.md

Embedding model, vector store, cross-language similarity, and the earlier 24B-with-manual against 70B-without result.

A2D

Study write-up (draft): Retrieval over Scale

V5/03 Artikkelit/Artikkeli 2 - RAG/artikkeli2_rag.tex

Cross-lingual retrieval probes, the Mistral Small 24B grounding comparison, and the three failure modes retrieval does not fix.

A3D

Study write-up (draft): Where the Pipeline Breaks

V5/03 Artikkelit/Artikkeli 3 - Putken virheanalyysi/artikkeli3_putki.tex

The three cross-stage failure modes, mitigations and the proposed propagation measurements.

SEL

First model comparison, 8 May 2026

V5/01 Dokumentaatio/Hynebot_V4_kuvaus.md, part 2

One question, five models, scored by eye. Kept because it was wrong.

Our own documents and code

STAT

V5 project status

V5/01 Dokumentaatio/projektin_tilanne.md, revision of 15 June 2026

The canonical running record of V5: decisions, lessons, open tasks, the 15 June correction.

PS

Progress summary, May–August 2026

V5/01 Dokumentaatio/Hynebot_V5_Progress_Summary_2026-08.docx

The source for the lessons in chapters 01, 03, 07, 08 and 11.

OV

V5 hardware notes

project notes, August–September 2026

Drivetrain parts on hand, speed limiting, scanner choice and safety-scanner provisions.

EMX

Motor-controller driver work, 25 August 2026

working notes on the EM-356A-SBL Modbus register map

The finding that the controllers are positioning drives, and the bench-test rule.

RISK

V5 safety memo and risk assessment

V5/05 GDPR ja turvallisuus/Hynebot_V5_turvallisuusmuistio.md

Version 1.0, 11 June 2026, awaiting approval. Ten hazards, technical safety functions, commissioning checklist.

GDPR

V5 privacy memo

V5/05 GDPR ja turvallisuus/Hynebot_V5_tietosuojamuistio.md

Version 1.0, 11 June 2026, awaiting the data protection officer.

DORNA

Dorna Arm 2 operating notes

V5/04 Robotiikka/DORNA_KAYTTOOHJE.md

Home position, parking pose, reliable command method.

ROUT

Router documentation

V5/01 Dokumentaatio/README_reititin.md

Single-model routing, the thinking switch, fallback.

PLAN

Research data plan

V5/02 Tutkimus/Suunnitelmat/tutkimusaineisto_suunnitelma.md

Question categories, acceptance criteria and red flags, evaluation method.

V4

V4 project summary and code

V1–V4/40 Hynebot V4/Dokumentaatio/Hynebot_V4_projektikooste_2026-05-26.md · Koodi

Hardware, services, network, control loop settings, display faces, lessons.

V4K

V4 description with V5 planning

V5/01 Dokumentaatio/Hynebot_V4_kuvaus.md

The spring 2026 decision to keep V4 and build V5 in parallel; the first V5 component plan.

V3

V3 wiring sheet and control code

V1–V4/30 Hynebot V3

ESP32, H-bridge, DC motors with Hall sensors, Raspberry Pi camera, browser video and control.

CTRL

V1/V2 control software

V1–V4/25 Ohjauskoodi V1–V2/hynebot-control (git)

Commits from March 2023 to February 2024.

V2R

V2 report and goals

V1–V4/20 Hynebot V2/Hynebot V2 Report.docx · Goals for V2.docx

14 August 2023. Aesthetic redesign, bearing housing fatigue, drive recommendation.

REQ

V1 requirements and wishes

V1–V4/10 Hynebot V1/Tables/Requirements and wishes.xlsx

Dated entries from August 2020 to early 2021, including the cost note of 13 October 2020.

SAFE1

V1 safety table

V1–V4/10 Hynebot V1/Tables/Safety_table.xlsx

Standards map and hazard mitigations, 2020–2021.

NUT

V1 specifications at a glance

V1–V4/00 Yleiset ja kytkentäkaaviot/Hynebot_info_in_a_nutshell.docx

Mass, height, footprint, drive, battery, computer, sensors, cost.

SWDEV

V1 software status, 21 June 2021

V1–V4/00 Yleiset ja kytkentäkaaviot/Hynebot_ohjelmistokehitys.docx

ROS 2 on WSL, driver gaps, controller USB problems.

Theses and conference papers

SEFI

Kasurinen, M., Ikävalko, M., Tuimala, L., Virkki-Hatakka, T. and Hyneman, J. Case: telepresence robot – virtual, but actively present teacher in a prototype laboratory

SEFI 48th Annual Conference, Enschede, 2020, concept papers, pp. 869–880 · V5/03 Artikkelit/Proceedings-DEF-nov-2020-kleiner.pdf

The V0 course project: requirements, design, the remote test of the sister robot in San Francisco, the design philosophy for presence, and the lessons on latency, testing and project style.

JUV

Juvonen, P. Mechanical Design of a Telepresence Robot for Instructional Use

Master's thesis, LUT University, 2021

The design of V1: requirements, safety standards, plywood construction, omnidirectional drive, cost.

ALW

Alwis Weerasinghe, S. Design of a Telepresence Robot: A Systematic Approach

Master's thesis, LUT University, 2024

Systematic redesign: two-wheel rear drive, MQTT, conferencing on a smart device.

HAK

Hakuli, M. Linux-pohjaisen robotin etäohjausjärjestelmän kehitys (Development of a Linux-based robot remote control system)

Master's thesis, LUT University, 2025

Remote-control requirements from interviews; server and socket.io architecture.

Published literature

MARCH

Marchisio, K., Dash, S., Chen, H., Aumiller, D., Üstün, A., Hooker, S. and Ruder, S. How Does Quantization Affect Multilingual LLMs?

Findings of EMNLP 2024, pp. 15928–15947

Quantization degrades lower-resourced and non-Latin-script languages most; evidence largely from translation and benchmarks at four bits.

BORG

Borgersen, K. A. and Goodwin, M. English K_Quantization of LLMs Does Not Disproportionately Diminish Multilingual Performance

arXiv 2503.03592, 2025

Finds no significant multilingual harm from K-quantization for one 70B model, in English, Norwegian and Malayalam, on a benchmark rather than generation.

CHANG

Chang, T.-Y., Zhang, M., Thomason, J. and Jia, R. Why Do Some Inputs Break Low-Bit LLM Quantization?

EMNLP 2025, pp. 3410–3429

Mechanism for input-dependent failure at three to four bits.

KIM

Kim, J., Ewer, E., Moon, T., Park, J. and Papailiopoulos, D. Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models

arXiv 2510.10964, 2025

The safe bit width depends on model scale; four bits is not universal.

BELCAK

Belcak, P., Heinrich, G., Diao, S., Fu, Y., Dong, X., Muralidharan, S., Lin, Y. C. and Molchanov, P. Small Language Models are the Future of Agentic AI

arXiv 2506.02153, 2025

The position that small, augmented models suffice for most agentic applications.

FULIU

Fu, X. and Liu, W. How Reliable is Multilingual LLM-as-a-Judge?

Findings of EMNLP 2025, pp. 11040–11053

Multilingual judge agreement: Fleiss κ around 0.3 on average, lower for some languages.

Scope and limits

The absolute numbers in this record apply to one robot lineage, one edge computer, one question set and one workshop. What is meant to transfer is set out in chapter 00, and the limits of the language-model results in chapter 10. Open items lists every contested figure found so far.

Licence, credit and how to cite

Text, tables, diagrams and data: Creative Commons Attribution 4.0 International (CC BY 4.0). You may copy, adapt and reuse them, including commercially, provided you credit the author, the J. Hyneman Center and this publication and indicate whether you changed anything. The photographs and renders are not covered by that licence and remain the property of their makers. Photographs credited to Teemu Leinonen are from LUT University's image bank and may be reused for communication purposes only, not for marketing, with the photographer credited.

Cite it as: Kasurinen, M. Building a robot for an open workshop: the Hynebot design record. JHC Platform Series no. 2. J. Hyneman Center, LUT University, revision 1.0, 2026. Cite the revision number; the numbers change between revisions.

Thanks

The course teams of 2019–2020 built the first robot, and Lauri Tuimala carried its software on. Perttu Juvonen, Sachintha Alwis Weerasinghe and Miro Hakuli wrote theses on it; Jake Cumens and Jethro McLean rebuilt it as V2 in the summer of 2023, and Robert Hämäläinen wrote its drive software in 2023–24. UPM Plywood supplied material for the first prototype and the LUT Climate Fund supported the work. Many others in and around the JHC team have contributed ideas, parts and time along the way. Three colleagues gave their time to score sixty answers blind.

Revision

Written by Marko Kasurinen, Head of Development at the J. Hyneman Center, LUT University. Published by the J. Hyneman Center, LUT University. Revision 0.1 of the English edition, 28 September 2026: first draft, assembled from the project archive with AI assistance and checked against the source documents by the author, who is responsible for the content, including its errors. Revision 0.2, 1 October 2026: second draft. Cover and photograph credits added; V2 authors named; V4 camera count corrected to three; V5 described as both telepresence robot and assistant throughout; the term “robot” clarified; Ollama and temperature explained; the reasons for the language-model studies and for publishing them here rewritten; the accessory-interface chapter withdrawn; the funding paragraph removed. Revision 0.3, 8 October 2026: third draft. Private correspondence is no longer used as a source, and the book no longer describes the plans of partner workshops; contributors are credited as members of the team rather than individually where they have not confirmed it. V5's body above the base, which was developed jointly, is no longer described; JHC's own work on the drivetrain, sensing, arm, brain, safety and privacy is kept. Revision 0.4, 9 October 2026: editorial revision for publication, covering structure, line editing, glosses for specialist terms, consistent terminology and number formatting, and checks of references, cross-references and navigation, with no change to any figure or finding. Revision 1.0, 9 October 2026: first published edition. The chapter on the robot generations now opens with their timeline and table; chapter titles were made parallel; the language-model table is ordered by Finnish accuracy; full author lists were added to the published literature; the remaining editorial notes were resolved. Open questions that belong to the project itself stay under Open items. New material goes in with a reference key, following the convention in chapter 00.