RTK Accuracy Trap: Can a Datasheet Predict Field Performance

AUTHOR: Zero Jiang | TITLE: Founder, Kalmix | READ: 15 min

TL;DR

  • A datasheet is the right first filter for an RTK receiver, but it cannot answer how the receiver behaves on your route.
  • Fixed Rate, RMS, and a clean-looking track answer different questions. None of them alone is a field-performance guarantee.
  • Ground-truth-aligned field testing shows what the receiver actually did across scenario changes, degraded windows, and solution states.
  • The selection question is how often the receiver reaches centimeter-level behavior, under which conditions, and how it degrades when the field stops matching the reference case.

No, a datasheet cannot predict RTK field performance by itself. It screens hardware capability; field performance depends on sky visibility, multipath, correction delivery, antenna setup, and how the receiver behaves during degraded windows.

Receiver selection often starts with a familiar comparison. Two RTK receivers both claim centimeter-level performance. The datasheet numbers look close enough that the choice seems to shift toward interface, price, power, packaging, or supplier trust. Then the team runs a real route: open road, tree cover, bridge edges, short correction interruptions, building edges, and a few sections where sky visibility changes faster than the receiver can comfortably recover.

That is where the evaluation gets harder. The same discussion now includes a datasheet spec, a Fixed Rate, one aggregate RMS, a trajectory screenshot, and maybe a vendor field test report. Those pieces of evidence are related, but they are not the same thing. Datasheet specs describe the hardware promise. Fix Type describes an algorithm state. RMS is a statistical expression of a test result. A track view is useful, but it is visual evidence, not an error distribution.

This article uses the GUIDE-K35E Ground Truth report as a reading case. It shows how an engineering team can read an RTK accuracy test before letting one headline number drive a purchasing or deployment decision.

What a Datasheet Number Can and Cannot Tell You

A datasheet is still the correct place to start. It tells an engineering team whether a receiver is even worth testing: supported bands and constellations, RTK support, output rate, interfaces, power, antenna requirements, mechanical constraints, environmental ratings, and nominal positioning performance. Without that first screen, a team cannot narrow the hardware set efficiently.

The trap appears when the headline accuracy number becomes detached from the conditions that define it. In mainstream RTK module datasheets, the small print often includes language such as "Depends on atmospheric conditions, baseline length, GNSS antenna, multipath conditions, satellite visibility and geometry." Test conditions may also reference a "1 km baseline" or "patch antennas with good ground planes." Those notes are not decorative. They define the world in which the accuracy statement was measured or expected.

Real deployments rarely hold those conditions steady. A robot may move from open sky into a tree canopy. A vehicle may pass under a bridge edge. A yard vehicle may spend part of its day near sheds, containers, fences, or partial overhead cover. The receiver is no longer operating inside the clean reference case implied by the datasheet.

That does not make the datasheet useless. It tells the team what to test next. If the accuracy claim depends on antenna setup, sky visibility, correction path, and geometry, then the field test needs to ask what happens when those conditions change.

For a deeper explanation of why CEP, RMS, R95, and other accuracy metrics are different statistical statements, see GNSS Accuracy Decoded. Here, the practical point is simpler: the same accuracy number can mean different things once it is attached to a route, a scenario, and a confidence level.

Why Real-World RTK Testing Is Still Necessary

Real deployment changes the GNSS problem. A receiver is not just solving satellite geometry in a lab. It is moving through sky visibility changes, multipath, signal discontinuity, correction coverage boundaries, vibration, antenna mounting effects, and host-system decisions about when to trust a position.

These changes do not require an extreme environment. A normal highway segment can produce a short degradation window when the vehicle passes a bridge, an overhead obstruction boundary, or the edge of a building. RTK may briefly lose lock, report Float, or show a short accuracy drop. Those few seconds are hard to predict from a datasheet.

A short highway obstruction window can push the solution state away from RTK Fixed. For a robot or vehicle, that brief low-confidence window is an operational event, not just a small change in route-average accuracy.

For a control system, the problem is not only the average performance over an entire route. The system consumes a continuous position stream. A short low-confidence window can affect speed control, localization weight, boundary behavior, and the decision to wait for re-lock before continuing a task.

Even when the receiver hardware is unchanged, the installation changes the result. Antenna position, vehicle body obstruction, cable routing, grounding, power noise, enclosure materials, motion dynamics, correction-source type, and correction-delivery quality all shape the output. A datasheet cannot fully enumerate those system-level variables.

That is why teams still need to test in environments close to their real deployment. The purpose is not to distrust the datasheet. The purpose is to place the receiver's datasheet promise back into the team's own task model, risk model, and control strategy. For robotics teams, this connects directly to trust-state design; the RTK GPS for Robotics guide covers the trust, degrade, and stop logic in more detail.

Why Field Testing Is Hard to Do Rigorously

Most teams already do some form of field testing. They drive a route, inspect the track, compare Fixed Rate, check whether the path stays near a map or marker, and decide whether the receiver feels stable enough. Those tests have value. They reveal integration problems, obvious degradation, bad antenna placement, and correction-path failures.

They are also easy to overread because GNSS inputs change with time. The receiver depends on the RF signal environment: sky visibility, multipath, NLOS, interference, satellite geometry, and correction delivery can all shift between passes. If a team runs one or two informal tests, it may confuse that day's conditions with a stable product difference.

Rigorous quantitative testing has to solve two problems at the same time: whether the test conditions are comparable, and what the true position is. Solving only one is not enough. Otherwise, the result can mix receiver behavior with satellite geometry, obstruction timing, correction state, or route execution.

LabSat and RF record-replay systems address signal-condition consistency. They allow a real RF environment to be recorded once and replayed for different receivers or firmware versions under the same signal conditions. They do not replace ground truth, but they make comparisons more consistent.

Ground truth equipment answers the other question: where was the vehicle or robot really located? Many robotics teams do not own tactical-grade FOG INS or an equivalent reference system. That equipment often requires a two-hundred-thousand to three-hundred-thousand-dollar level of investment. Without a high-grade reference, it is hard to turn "the track looks okay" into a rigorous horizontal error distribution.

Evidence Layers

A quick on-site drive can reveal integration problems. A ground-truth-aligned report can quantify error distribution. RF record/replay can improve repeatability across receivers or firmware versions. These are different evidence layers, not interchangeable proof.

This applies to Kalmix as well. SCOUT PRO uses a GUIDE-K35E-related positioning path, but a module-level test with a survey-grade antenna is not the same as finished-product performance. Enclosure, antenna, mounting, power, firmware configuration, thermal behavior, and the real application environment still matter. A mature report should state those boundaries instead of asking the reader to infer them.

The Fix Rate Trap: Fixed Is Not an Accuracy Certificate

A higher Fixed Rate is often misinterpreted as higher accuracy. Fixed Rate is useful, but it only tells part of the story. It tells you how often the receiver reported a fixed solution under that route and correction setup.

Fixed and Float are solution states, not accuracy bins. Fixed means the receiver's algorithm has accepted an integer ambiguity solution. It often corresponds to centimeter-level performance, especially under strong conditions, but it does not prove that every Fixed epoch met a specific error threshold. Float means the ambiguity is not fixed. Its error may range from useful to several meters depending on the environment and application threshold.

The difference matters for machines. High-confidence wrong data can be harder to handle than honestly low-confidence data. A robot can down-weight Float, slow down, or wait. It is much harder to recover when a position stream looks trustworthy but does not match reality.

For accuracy evaluation, solution state has to be checked against measured error. In the K35E route test, the Fix / Float cumulative distribution below compares receiver-reported solution state with ground-truth-aligned horizontal error. It shows what the receiver reported and how far those reported states were from the reference trajectory.

Interactive CDF: RTK Fixed vs RTK Float Horizontal Error

Data source: GT-01-K35E. X-axis uses log scale so centimeter-level Fixed behavior and meter-level Float behavior can be read in the same chart.

RTK Fixed RTK Float

Fixed reading: In this K35E route test, about 90% of Fixed epochs are within 3 cm, and 99.78% are within 10 cm. That is evidence from measured error distribution, not from the Fixed label alone.

Float reading: In this K35E route test, Float is not a stable accuracy grade. About 80% of Float epochs are within 5 m, with a visible tail beyond that threshold.

The chart here is included to demonstrate how to read solution-state distributions, not to turn one route into a universal product guarantee.

The Fixed curve confirms centimeter-level precision under these test conditions. That does not make Fixed a universal accuracy certificate. It means the K35E Fixed state aligned very well with ground truth in this route test.

The Float curve exposes unpredictable meter-level degradation. Float is not one precision level. It can sit near a task's tolerance in one moment and become unusable in another. A low-speed campus vehicle, a mower working near a boundary, and a mapping app collecting points may all set different thresholds. The receiver state matters, but the measured error distribution decides whether that state is usable.

For the RTK mechanism behind these states, see The RTK Trick. For evaluation, the important rule is: Fixed Rate shows how often the receiver reported a fixed solution. The CDF shows the measured error behind those reported states.

Read by Scenario, Not by One Aggregate Number

An aggregate metric cannot be read without scenario context. RMS is a statistical expression of test results. It can mix open road, short obstruction, sustained overhead obstruction, loss-of-lock, and re-convergence windows into one number.

Because RMS squares the error before averaging, larger errors from multipath, blockage, and re-convergence windows can dominate the aggregate statistic. A mixed-route RMS therefore says as much about scenario composition as it says about receiver behavior. That does not make RMS useless. It means the reader must ask what kinds of scenes are inside the statistic.

The reading order should be simple: first identify the scenarios in the test, then read each scenario's CEP50, R95, RMS, solution mix, CDF, and tail behavior. Scenario statistics are not an appendix to a field report. They are the way aggregate metrics recover engineering meaning.

The K35E test data is useful as a case because it allows the reader to move beyond one mixed-route headline. In this article, the incremental read is to isolate a severe overhead-obstruction window and examine it as its own scenario instead of leaving it buried inside a mixed-route summary.

Do not read an aggregate number without scenario context

RMS, R95, and Fixed Rate only become useful when you know how much open sky, obstruction, transition, and re-convergence time the test contains.

Incremental Scenario Evidence: What Overhead Obstruction Reveals

The published report does not name Overhead Obstruction as a separate scenario. For this article, we isolate that degraded window as an incremental analysis: the vehicle repeatedly passes under elevated roadway or bridge structures, where the top-side satellite view is heavily blocked and signal continuity is compromised.

In this K35E route test, the Overhead Obstruction segment lasts about 998 seconds across 7.481 km and 9,981 epochs. RTK Float accounts for 97.85% of the segment. Horizontal CEP50 is about 2.86 m, R95 is about 7.42 m, and RMS is about 3.76 m.

Duration 998 s 7.481 km segment
Epochs 9,981 10 Hz route data
RTK Float 97.85% Fixed only 2.15%
Horizontal R95 7.42 m CEP50 2.86 m

These statistics represent a single route snapshot to illustrate evaluation methods. For the definitive multi-scenario error distribution, refer to the full K35E Ground Truth Report.

Full Overhead Obstruction scene. Repeated passes under overhead roadway structures link the segment statistics to sustained top-side sky blockage.

This is not a hardware failure; the receiver is measuring a different GNSS problem. When the top of the sky is persistently compromised, a receiver that performs tightly in Fixed conditions can spend most of the segment in Float and move into meter-level error.

Overhead structures do more than remove satellites from view. They also increase multipath and NLOS risk, which contaminates measurements and makes integer ambiguity validation harder. Dropping to Float is often the receiver avoiding a false fixed solution under corrupted geometry.

This segment also explains why solution-state distribution should be read by scenario. The full-route CDF contains 10,367 RTK Float epochs; this Overhead Obstruction segment accounts for roughly 94% of them in this test. The route does not degrade evenly, so the report should not be read as if it does.

This behavior matters for robots, low-speed vehicles, yard vehicles, campus vehicles, and field machines that must operate near bridges, sheds, tree canopies, building edges, or other weak-sky environments. In this test, the contrast is also instructive: a high-rise urban canyon is not automatically harder than a lower, continuous overhead obstruction if the latter blocks the sky more persistently.

How Engineering Teams Should Evaluate RTK Receivers

A practical RTK evaluation does not replace datasheets with field videos, or replace field reports with one route of internal testing. It builds a chain of evidence.

In practice, that runs through four steps: screen the hardware, design the route, log the right diagnostics, and report statistics at the scenario level.

Start with the datasheet. Use it to screen bands, constellations, RTK support, output rate, interface, antenna requirements, power, mechanics, and environmental constraints. That step decides which receivers deserve field time.

Then design the route. Separate open sky, typical route, weak coverage, transition / re-lock windows, and the specific obstruction cases your deployment will actually see. A route that only shows the cleanest behavior is not a deployment test. A route that only punishes the receiver is not a fair baseline either.

During the test, log more than coordinates. Fix Type, correction age, satellite count, DOP or signal metrics, receiver-reported uncertainty, timestamps, state transitions, and diagnostic logs are what allow a team to explain why performance changed.

Finally, report statistics at the right level: scenario-level CEP50 / R95 / RMS, solution mix, accuracy CDF, Fixed-state error distribution, Float-state error distribution, re-convergence windows, and data completeness. If ground truth is not available, say so. If RF conditions were not repeatable, say so. Evidence is still useful when its boundary is clear.

The resulting evaluation hierarchy maps directly to deployment risk:

Evidence type What it can tell you What it cannot prove alone
Datasheet review Whether the hardware is worth testing. Real route accuracy.
Simple field drive Integration issues and obvious degradation. Quantitative error distribution.
Repeatability test Whether the system returns to similar paths under similar conditions. Absolute accuracy versus ground truth.
Ground-truth-aligned test Error distribution, tails, solution-state behavior, and scenario behavior. Performance in every deployment.
Your own deployment test Fit for your actual robot, route, antenna, mounting, and workflow. A general product guarantee for every other site.

This framework can become the starting point for an RFP, vendor evaluation, or internal test plan. A defensible RTK evaluation is not one perfect number. It is a chain of evidence, from datasheet to controlled report to the operating environment that will carry the risk.

Conclusion: Specs Start the Conversation; Field Evidence Decides the Risk

Datasheets help you decide what is worth testing. Ground-truth-aligned reports help you understand observed performance, degradation behavior, and statistical tails under a specific setup. Your own deployment testing decides whether the receiver fits your machine, site, workflow, and risk threshold.

Kalmix is not asking teams to ignore datasheets or blindly trust a vendor report. The practical standard is simpler: read every datasheet number as a conditional engineering statement, then read every field result with its scenario and statistical meaning. Finally, test in the environment where the receiver will actually work.

For future reports and the test-library context, visit Kalmix Ground Truth field tests.

Key Takeaway

Do not compare RTK receivers by one number. Compare the evidence layer: spec, solution state, distribution, scenario, and your own deployment route.

Frequently Asked Questions

Does RTK Fixed always mean centimeter accuracy?

No. RTK Fixed means the receiver has accepted an integer ambiguity solution; it usually corresponds to centimeter-level accuracy under good conditions, but it is not a standalone accuracy guarantee. It is a state label, not a measured error percentile. For evaluation, Fix Type should be read together with ground-truth-aligned error distribution, correction age, satellite geometry, and scenario context.

Why can RMS look worse than typical RTK accuracy in a field test?

RMS is sensitive to larger errors and degraded segments. A test that includes more obstruction, weak coverage, loss-of-lock, or re-convergence time can show a larger RMS even when many points remain accurate. The same receiver can therefore show different RMS results when route composition changes. That is why RMS should be read with scenario composition, not as a standalone receiver label.

What is the best way to evaluate RTK receiver accuracy?

Start with the datasheet to screen hardware capability. Then read a ground-truth-aligned field report for measured error distribution, solution-state behavior, and scenario-level statistics. Finally, run your own deployment test with the intended antenna, mounting, correction source, route, and operating thresholds. This keeps vendor evidence useful without treating it as a deployment guarantee.

Can I evaluate an RTK receiver by Fixed Rate alone?

Fixed Rate is useful, but it is not enough. It shows how often the receiver reported a fixed solution. It does not show the measured error behind those states, how Float behaved, how large the tail errors were, or how the receiver recovered after obstruction or correction interruptions. A CDF or error histogram is needed to read the distribution behind the state label.

How accurate is RTK under bridges or overhead obstruction?

RTK accuracy can drop from centimeter-level behavior to meter-level error under sustained overhead obstruction. The exact result depends on sky blockage, multipath, antenna placement, correction continuity, and receiver recovery behavior. This is why obstruction should be evaluated with its own scenario statistics rather than folded into one mixed-route number.

You Might Also Like

Zero Jiang, Founder of Kalmix

Zero Jiang

Founder, Kalmix

Dedicated to making high-precision GNSS positioning accessible and reliable for global developers building autonomous systems and robust RTK hardware.

Evaluate the GUIDE module path, then compare measured behavior across Ground Truth reports.

Back to blog