Time
A second is a second, I get it. But when someone says a sensor has 150 µs latency, it sounds tiny until you compare it to a 33,000 µs camera frame. Fourteen orders of magnitude live in a robotics stack. The time axis is interesting because things happen at different rates, and you have to know which clock is governing your problem.
For a robot moving at 1 m/s, a 33 ms camera frame equals 3.3 cm of blindness. At 2 m/s it's 6.6 cm. A useful change in perspective for me came when I stopped thinking in milliseconds and equated it to distance unseen between vision signals.
The MIT Cheetah's foot is in the air for ~85 ms — fewer than three 30 fps frames. Your control loop has this long to decide what to do when it lands. Human voluntary reaction time floors at ~250 ms. It is not a coincidence that this is where network teleoperation becomes unstable.
Frequency
Hertz is just events per second. 1 Hz means once per second. 1 kHz is a thousand times per second. The reciprocal relationship is most useful in real-time systems: period (s) = 1 / frequency (Hz). A 1 kHz control loop runs every 1/1000 = 1 ms. A 30 fps camera produces a frame every 1/30 ≈ 33 ms. A 4 GHz CPU completes a clock cycle every 0.25 ns. These are the same facts written two ways.
A common rule of thumb is that each control layer must run ~10× faster than the layer above it. This is why robotic stacks have 1 kHz current loops under 100 Hz position controllers under 10 Hz planners. For example, a VLA at 5–50 Hz commands a manipulator executing at 50 Hz with action chunking, so the policy recomputes every 73 ms but the robot fully executes the pre-predicted chunk between recomputes. Slow policy, fast execution.
Bandwidth
Bandwidth is information per unit time. The intuitive picture is a pipe. A larger pipe moves more water per second, and the same is true of memory buses or network links.
Bandwidth matters in robotics because you cannot do math on bits that haven't arrived yet. If your model has 70 GB of weights and your memory bandwidth is 3.3 TB/s, the minimum time to read every weight once is 70/3300 = 21 ms. A useful intuition is HBM bandwidth on a top-end GPU or TPU is roughly the same as moving a Christopher Nolan 4K movie every second. It sounds absurd until you realize a 70B-parameter model wants to be read end-to-end every few milliseconds during inference.
FLOPS
FLOPS = floating-point operations per second. One FLOP is an add or a multiply. Like Hz, FLOPS are a rate. They tell you how quickly a processor can do math.
Precision matters here. A FLOP at FP32 and a FLOP at FP4 are not the same physical operation. FP4 works on 8× less data than FP32 and usually 10× less energy, so when NVIDIA says Blackwell does 20 PFLOPS, they could mean FP4 with sparsity tricks.
Energy & Power
This is the most important pair of units in the post. Energy is the amount of work done measured in joules (J) or watt-hours (Wh). Power is the rate at which energy is delivered in watts (W) = 1 joule per second. Energy = Power x Time. A 100 W lightbulb left on for 1 second consumes 100 J. Left on for 1 hour, it consumes 360 kJ. A DRAM read costs 200× an FP multiply — on the same chip. Moving bits costs more than computing on them. Every architecture bets on this ordering.
I prefer listing arithmetic operations in picojoules (pJ) because energy is what matters. A DRAM read costs 200× more than a floating-point multiply on the same chip. Accelerators do trillions of FLOP/s, so picojoules-per-FLOP multiplied by FLOP/s gives you watts of power consumed. A 3.7 pJ multiply × 1 trillion = 3.7 watts. This tiny per-operation number done over and over is what determines whether your robot needs a fan.
Latency vs Throughput
Latency is how long one thing takes. How long until the first token comes back? How long from sensor reading to motor command? Throughput is how many things happen per unit time. How many tokens per second can the GPU emit? How many camera frames does this pipeline process per second? You can have one without the other: NVLink72 is this idea exactly. Enormous aggregate token throughput but the individual token latency is unchanged.
| Quantity | Unit | Robotics example |
|---|---|---|
| Latency | ms, µs, ns | Sensor reading to motor command; time to first token |
| Throughput | Hz, FPS, tok/s, GB/s | Control-loop rate; camera frames per second; tokens per second |
For AI serving, you usually choose the opposite trade-off. Batching improves throughput at the cost of latency. If you wait until 64 inference requests arrive before processing, you spread the cost of reading the weights across all 64, but the first request waited for the other 63 to show up.
For robots, latency is non-negotiable. A controller that processes 10,000 sensor readings per second on average but occasionally takes 50 ms to respond to one of them will break.
I've been working on the Intrinsic AI for Industry Challenge, where the goal is to automate cable and wiring tasks. 2.3 kWh batteries, 90 ms VLA inference time, 5 Hz clock cycle, 640 pJ DRAM read, 250 ms human teleop reaction time. A working policy is a stack of constraints in so many units. Early in my robotics journey I have discovered robotics is a timing problem. Every successful design is, at its core, a back-of-the-envelope reconciliation between control loops and clock cycles and latency and inference times. Once you get this right, then it's an energy problem.
Sensors
| Sensor | Res / rate | Latency | Bandwidth |
|---|---|---|---|
| RGB camera | 1080p @ 30 fps | 33 ms/frame | 187 MB/s raw |
| Ouster OS1-128 Rev7 LiDAR | 5.2 M pts/s @ 10–20 Hz | 50–100 ms | ~50 MB/s |
| VectorNav VN-100 IMU | IMU/AHRS | <1 ms | <1 MB/s |
| Renishaw RESOLUTE Encoder | 32-bit optical | <10 µs | — |
| ATI Mini40 Force/Torque | 6-DOF, up to 7 kHz | ~140 µs | — |
The Control Hierarchy
Energy & Batteries
A human burns through these average humanoid evergy stores in about 12 hours.
| Robot | Battery | Runtime | Implied avg power |
|---|---|---|---|
| Boston Dynamics Spot | 580 Wh | 90 min | ~390 W |
| Unitree H1 | 864 Wh | 1.5–2 hr | ~500 W |
| Tesla Optimus | 2.3 kWh | Workday (allegedly) | — |
| Figure 03 | 2.3 kWh | 5 hr | ~455 W |
Worked example
A Figure 03 robot 2.3 kWh pack can support a five-hour shift at about 460 W average. An eight-hour target forces the whole robot under 288 W, so a 70 W edge computer plus sensors is feasible only if walking and manipulation stay light.
Edge Compute
| Module | Performance | Memory BW | Power |
|---|---|---|---|
| Jetson Orin AGX | 275 TFLOPS | 205 GB/s | 15–60 W |
| Jetson Thor (T5000) | 2,070 TFLOPS (FP4) | 273 GB/s | 0–130 W |
Thor is the first edge SoC where running a 7B-parameter VLA above 10 Hz is comfortable. This matters: at <10 Hz a policy must chunk aggressively to maintain 50 Hz arm execution. Orin and Thor both ship hardware video encoders, vision accelerators, and ISP pipelines to offload video encoding to dedicated hardware blocks.
Communication Buses
| Bus | Bandwidth | Jitter / cycle | Use case |
|---|---|---|---|
| ROS 2 DDS | — | 50 µs–10 ms | Vision, planning, and non-real-time coordination |
| EtherCAT | 100 Mbps–1 Gbps | <1 µs | Deterministic joint control and servo synchronization |
| GMSL2 | 6 Gbps | <100 µs | Perception bandwidth |
36-DOF actuator state at 1 kHz × 32 bytes = 1.2 MB/s. Four 4K cameras raw = 6 GB/s. These differ by more than 5,000×. Design the buses accordingly — one is trivially served, one requires on-sensor encode to be feasible at all.
VLAs & The AI Stack
| Model | Params | Hardware | Inference | Exec rate |
|---|---|---|---|---|
| RT-1 | 35 M | TPU | 100 ms | 3 Hz |
| OpenVLA | 7 B | RTX 4090 BF16 | 167 ms | 6 Hz |
| π₀ (Chunk 50) | 3.3 B | RTX 4090 | 73 ms/chunk | 0 Hz exec |
| π₀.₆ | 5 B | H100 | 63 ms/chunk | 50 Hz exec |
| Helix S2 (Figure) | 7 B | Embedded GPU | — | 7–9 Hz |
| GR00T N1 (NVIDIA) | — | NVIDIA L40 | — | 10 Hz (S2) |
Action chunking: π₀ predicts 50 future actions per 73 ms inference. The robot executes the chunk at 50 Hz and the policy recomputes in parallel. Original ACT imitation success jumps from 1% at chunk=1 to 44% at chunk=100. The bottleneck isn't GPUs. Figure, Pi, and Tesla all internalized this in 2024–25: teleop data collection costs 10–100× more than the compute to train on it.
Robot Training Data Layer
Offline robot learning is governed by episodes, timestamps, video codecs, samples/s, and GPU starvation.
| Layer | Unit to think in | Hidden tax |
|---|---|---|
| Recording | GB/hr, dropped frames, clock drift | Bad timestamps and missing streams |
| Compression | GOP length, keyframe interval | Smaller files can make random frame access expensive |
| Sample construction | frames/sample, columns/sample, history window | A training sample is a time-aligned slice, not one row |
| Dataloader | samples/s, GPU utilization | Remote fetch and decode stalls waste accelerator time |
| Curation | episode weights, task mix, failure filters | Slow exports make dataset iteration feel like pipeline work |
Worked example: video training sample
Compression saves storage, but long group of pictures (GOP) video turns random access into seek-and-decode work. With a GOP of 30, fetching one arbitrary frame can require decoding about 15 frames on average before yielding one usable image.
Teleop Latency
Biological Comparisons
The most successful robot architectures will probably reverse-engineer what evolution spent 500M years refining.
| Biological system | Value | Robot analogue |
|---|---|---|
| Nerve Conduction | 70–120 m/s | EtherCAT propagation at ~200 m/µs |
| Synaptic Reflex | 25–35 ms | Motor current loop period: 1 ms |
| Voluntary Visual Reaction | 200–250 ms | VLA inference: 73–167 ms |
| Free Viewing | ~5 Hz | Camera capture rate: 30–90 fps |
| Resting Metabolic Rate | 80–100 W | Jetson Orin at 30 W ≈ ⅓ of resting human |
| Sustained Athletic Peak | 400–500 W | Walking humanoid average: 150–250 W |
| Skeletal Muscle Efficiency | 18–26% | Equivalent to an ICE; BLDC: 85–95% |
Humans run a slow conscious loop (~5 Hz) over a fast reflex layer (~25 ms).