Why Robots Look Jerky at Low Control Frequencies

Quick note on scope: When getting into robot learning from a background in machine learning, not robotics, I spent some time thinking about why some robot movements are jerky and others are not. This blog post is a summary of simulation experiments that helped me figure it out without touching a textbook. So don’t expect a complete, textbook-style coverage of robot control, instead this is meant as a small shortcut for you if you are new to robot learning.

What is the control frequency?

We control the robot by submitting target states for its angles to it. The control frequency means ‘how many times per second do we submit a new target state to the robot’. E.g. a control frequency of 50 Hz means we submit a new target action to the robot 50 times per second, i.e. every 20 ms.

In our setting, we assume that changing the control frequency means only changing the time that passes between submitting two actions to the robot, not changing the sequence of target states itself (i.e. we always submit the same sequence of degrees to a joint, just spaced out differently over time). This is realistic because in robot learning the policy that controls the robot is a function of the robot’s state (in the simplest case only its joint angles) and not of the wall-clock time. In theory (up to backlash, overshoot etc.), we should get the same actions from the policy whether we wait 100ms or 20 seconds as long as the robot had reached its previous target state, i.e. it does not move anymore while we wait. In practice, one might encounter a reduced control frequency if the policy inference is too slow or if one wants to slow down the rollout (more below on how to do this without introducing jerkiness). Note that in this setup the higher the frequency the faster the robot moves.

When does a robot look jerky to us?

If it changes its velocity a lot within a small amount of time. The movement looks completely smooth if the speed is constant throughout. If it moves very fast, then very slow it looks jerky (cf. Videos 1). We can characterize it as \(\text{jerkiness} = \frac{ \max_k(v) - \min_k(v)}{k}\), where \(v\) is the velocity and \(\max_k\) and \(\min_k\) refer to taking the max/min across a sliding window with size \(k\) timesteps respectively. This is a simple proxy for how jerky the motion looks (in robotics, smoothness is formally measured via jerk, the derivative of acceleration).

Video 1a: 10 Hz control frequency.
Video 1b: 50 Hz control frequency.

The frequency at which we submit actions to a robot can influence both the maximum and minimum velocity the robot reaches. Now, we’ll figure out how exactly.

How is the velocity of a robot determined?

In general terms, the robot turns ‘distance to target’ into ‘speed right now’. This gives an exponential decrease in the distance to the target, so the robot moves fast at first and then slow. If we submit a single action to the robot like in Video 2, the velocity over time of the shoulder pan joint looks like Figure 2:

Figure 2: Velocity and position of the shoulder pan joint after submitting one target state at t = 0. The robot accelerates, overshoots the target and oscillates before settling.
Video 2: A single target state submitted to the robot.

If we submit such an action to the robot once per second (1 Hz control frequency), we can see how this pattern simply repeats (Figure 3, Video 3):

Figure 3: Velocity (left) and position (right) at 1 Hz control frequency. Each action triggers the same spike-and-settle pattern as in Figure 2, followed by a long standstill.
Video 3: 1 Hz control frequency (Figure 3).

Now, depending on how quickly we submit the next action, only parts of this pattern realize and the observed velocity pattern changes (Figure 4). This is the key to the smoothness of higher frequency control.

Figure 4: Left: velocity at 1 Hz vs. 10 Hz over time. At 10 Hz the settling jitter is cut off, but the peaks stay the same. Right: velocity at 10 Hz vs. 50 Hz over the action index. At 50 Hz the velocity drops much less between actions.

If we submit the actions 10 times per second (10 Hz), we see in the left plot of Figure 4 that only the jitter around 0 velocity where the robot is settling in the new state is cut off by submitting the next action. The minimum and maximum velocity stay the same.

If we now submit the next actions even closer in time, the next action comes in early during the slowing down phase. If that action points in the same direction it immediately results in a speed up towards that new target, before the robot even gets the chance to slow down much.

At 50 Hz (right plot of Figure 4) we see that while the peak velocity stays the same, this results in way less slowing down. Hence, in any small time frame the minimum velocity is much closer to the maximum velocity than at 10 Hz and thus the robot movement looks much less jerky. (Note that for the right plot of Figure 4, we used the action index as the x-axis to align the velocity measurement with where on the circular trajectory the robot arm is at any given moment since that also influences the velocity level. For the 1 Hz vs. 10 Hz comparison this does not play a role because the 1 Hz script makes very little progress on the trajectory anyway.)

The movement is visibly smoother at 50 Hz than at 10 Hz (cf. the videos in the beginning: Videos 1)).

What happens if we increase the control frequency far beyond 50 Hz?

Since we saw that the robot movement becomes smoother the higher the control frequency, one might wonder what happens if we keep increasing the frequency (assuming our policy inference could keep up). Unsurprisingly, we will run into problems eventually. If we look again at the velocity pattern, we see that if we submit the actions very closely after each other, we will submit the next action while the robot was still accelerating towards the previous target. If the robot is still accelerating, this means that it had not come close to the previous target state yet. This can accumulate a lag because the targets now move faster than the robot can follow. Hence, the robot strays from the trajectory that we submit to it (Video 5).

The circular trajectory in Figure 5 demonstrates this well: We can see that the robot’s trajectory never catches up to the proper submitted target states. Additionally, you can see it taking shortcuts since it never reaches the max and min degrees that are submitted to it as target positions.

Figure 5: Joint position over the action index at 50 Hz and 200 Hz, with the submitted target positions dashed. At 200 Hz the robot lags behind the targets and never reaches their full amplitude.
Video 5: 200 Hz control frequency. The arm lags behind the targets and cuts the circle short.

Implications for training a neural network to control a robot

We just saw that a robot moves smoothly if the time we give the robot to reach its next target state is just enough to reach it. Hence, at higher control frequencies the distance (in terms of degrees) to the target states should be smaller than at lower control frequencies. This is fundamentally relevant if we want a neural network to control the robot smoothly. The neural network needs to spend substantial effort learning how far away an action it submits may be from where the robot currently is such that the robot still moves smoothly.

This directly explains why if we train a policy with behavioral cloning the frequency of data collection from teleoperation should be the same as the control frequency at which the policy is then controlling the robot. The teleoperator will move in a way that makes the robot move smoothly. In the dataset the smoothness does not show up directly, the policy simply learns to replicate how far (in terms of degrees) actions should be spaced apart. If you then, e.g. decrease the control frequency at deployment time (e.g. could happen if the policy inference is too slow), the policy still outputs actions that are just as far away from each other as in the training data but the robot gets there too early and slows down a lot, hence jerky movement.

Cheat Code: Action Interpolation

The idea is to artificially split an action into multiple substeps that we then submit to the robot one after the other. Instead of submitting the ‘real’ next target state directly to the robot, we compute waypoints between where the robot is now and the target state, interpolating linearly. Say, a joint is currently at 0 degrees and the next target state is to have the joint at 9 degrees. With an interpolation multiplier of 3, we would submit the following sequence of target states to the robot: [3, 6, 9]. Thus the maximum distance to target the robot ever sees is 3 instead of 9 and so the maximum velocity is lower too. For a simple pan, we can immediately see in Figure 6 and by comparing Video 6a with Video 6b how the interpolation decreases the magnitude of the velocity and delays the arrival at the target state.

Figure 6: Velocity (left) and position (right) of a single pan with interpolation multipliers 1, 5 and 40. Higher multipliers lower the peak velocity and delay the arrival at the final target.
Video 6a: Pan without interpolation (multiplier 1, same as Video 2).
Video 6b: Pan with interpolation multiplier 40.

Usually, action interpolation is done to increase the effective control frequency while keeping the speed of the robot the same. To achieve this one would still submit the non-interpolated actions at the original control frequency and then spread the interpolated actions evenly in the interval of time between two non-interpolated actions. In our case of the simple pan, the two cases are equivalent because there is only one non-interpolated action, hence there is no notion of control frequency for the non-interpolated action in this specific setting.

How can we slow down the robot without making it jerky?

Sometimes after training a new policy to control the robot, it can be scary to try it out on a real robot for the first time since it can break real things. To make it safer people like to artificially reduce the velocity at which the policy moves. The question is then how to do this without altering the policy’s planned trajectory (too much).

Action interpolation gives one way to do this. We simply decrease the control frequency at which we roll out the policy by a factor of \(n\) and then use \(n\) as our action interpolation multiplier. For each action that the policy produces, we compute \((n-1)\) interpolated actions that lead linearly to the policy’s action. Like this the robot moves at the speed of a lower control frequency but with the smoothness of the original control frequency. In Figure 7, the slowed-down rollout follows virtually the same positions as the original one, at much lower and less fluctuating velocity (Video 7a vs. Video 7b).

Figure 7: 10 Hz with interpolation multiplier 10 vs. 50 Hz without interpolation, over the index of non-interpolated actions. The positions (right) are virtually identical, while the velocity (left) is lower and fluctuates less.
Video 7a: 10 Hz with interpolation multiplier 10 (slowed down).
Video 7b: 50 Hz without interpolation.

Appendix

Querying a Theoretical Trajectory (Thought Experiment for Better Understanding)

Unlike the rest of this post, here we assume that we have an underlying trajectory that is fixed per period of time. E.g. a function that takes in the current time and gives out target angles for each robot joint such that once per second, the robot gripper moves in a full circle. We can then query this function whenever we like and assume we immediately pass the target state the function gives as a response to that query to the robot. Thus, in this case, the control frequency defines how long we wait between each query. If we reduce the frequency, we wait longer before asking again where the robot should move to.

For movements in a straight line this case maps nicely onto the simple ‘remaining distance proportional to current speed’. If we decrease the frequency at which we query and submit actions, the distance between these actions increases because the underlying trajectory keeps moving while we wait. Because of the larger maximum distance the maximum velocity increases and thus the difference between maximum and minimum velocity increases, hence more jerkiness. Note that if the underlying trajectory were in fact a circle, the realized movement of the robot strays further from the theoretical circle the further we decrease the control frequency (compare Video 8a with Video 8b). As an extreme: if we were to reduce the control frequency to 2 Hz and the circle completes once every second, this means we only ever get the same two alternating target states from the circle. Since the robot always moves in a straight line to its target state it would not move in a circle but just in a straight line between two points. If we were to query the circle at 1 Hz, the robot would not move at all after it reaches that point for the first time.

Video 8a: Querying the circle at a higher frequency.
Video 8b: Querying the circle at a low frequency. The robot waits, then moves in large steps.