Flight test results · Cessna 172 · Simulation

Vision-only approach guidance

An aircraft that finds its runway and flies the approach with no satellite navigation and no radio aid, with the decision-making on an embedded module. The same 162 deliberately wrong starts have now been flown twice, by two versions of the system. It lines up to within 4 m from anywhere. From 1.5 nm it lands on the threshold; from 2.5 nm it still lands 160 m short. Every number is here, including the ones that look bad.

Joe Tan Singapore Personal project · personal hardware Updated
11–52 m
Position error where the approach began, with no GPS
After 107 and 172 nm
162 ×2
The same starts flown by two versions of the system
3 runway ends · 54 starts each
4.2 m
Median error from the centreline at touchdown, from any starting offset
Was 16.6 m
59%
Of touchdowns still more than 100 m short of the aim point
Was 68% · the open problem
NAVnavigation

No satellites, no beacons, no beam

Jamming and spoofing have made satellite navigation the softest part of an aircraft's stack, and the approach is where losing it hurts most. This project asks a narrow, testable question: with the GPS gone and no ground aid at the field, can an aircraft still find a runway at the far end of a 172 nm leg and fly itself onto it?

The aircraft is a Cessna 172 — a four-seat single-engine light aeroplane, and the most ordinary aircraft it is possible to choose. That is deliberate: nothing here rests on an exotic airframe, a bespoke sensor or a research platform nobody else can get hold of.

The navigation half works. Across two routes — 107 nm and 172 nm over real terrain, in wind that changes along the track, with gusts — the aircraft's own estimate of where it was sat 11 to 52 m from the truth by the time the approach began, with no false position fixes in the verification passes. That is close enough to hand a runway to the approach, and it is the number everything below depends on.

The approach half is where the results get interesting, and it has just moved. An earlier version of the approach system lined up well and would not land — it arrived low, fast and flat, and almost never flared. The version reported here flares on four touchdowns in five. Side to side is now solved. What is left is a single, much narrower problem: how far out the approach begins. From close in it lands on the threshold. From further out it still lands short, and by almost exactly as much as it used to.

Both versions were flown on the identical 162 starts, scored the same way, so every comparison on this page is like for like. The older figures are kept alongside the new ones rather than replaced.

FILMin flight

Watch it fly

Rendered from the recorded trajectory, with the runway the system is tracking marked, live distance, height, speed and sink, and a touchdown card measured against the runway at the end. This clip comes from the wider runway campaign rather than the three-airfield matrix the figures on this page are drawn from, and it shows the earlier of the two versions compared above — it is not a recording of the results on this page.

ENVthe envelope

From where can it actually recover?

A demo answers “does it work?” An envelope answers the question an operator actually has: from how far wrong can this thing recover, and at what point does it give up? So the system was started deliberately wrong — 162 times, one error at a time, at three ranges from the runway, from up to 400 m to either side and up to 300 ft off the correct height. Then the whole matrix was flown again, unchanged, by the newer version. The charts below are the newer one.

Started off to the sideGo-arounds

centred0 / 9
50 m0 / 18
100 m0 / 18
150 m2 / 18
200 m1 / 18
300 m4 / 18
400 m7 / 18

Started off the correct heightGo-arounds

−200 ft2 / 9
−100 ft0 / 9
+100 ft1 / 9
+200 ft0 / 9
+300 ft0 / 9

Distance out at the startGo-arounds

1.5 nm13 / 54
2.0 nm4 / 54
2.5 nm0 / 54
Recovers Outside the envelope — gives up 40% of the time or more. Nothing now reaches it All three charts share one scale
The measured approach envelope. A go-around is counted as a failure to recover here, because each start was given a single attempt. Total go-arounds fell from 23 to 17 between the two versions. The envelope now reads: offsets to 100 m are absorbed completely — not one go-around — and the strain only shows at 300 m and beyond, with 400 m sitting just inside the 40% line. Height errors have almost stopped mattering; only arriving 200 ft low still costs anything, and 300 ft high, which used to be the worst case, now costs nothing. The range chart is the one to read twice, because it inverts: the closest starts produce the most go-arounds and the best landings, while from 2.5 nm it never gives up at all and lands 163 m short instead. Giving up early and landing short are the two ways this system fails, and distance decides which one you get. An earlier campaign also tested 1.0 nm starts and found them unusable — about half were abandoned before any intercept began — so that range was dropped and is not in these charts. Each cell is a single flight per runway, so these are trends, not rates.
LDGtouchdown

What changed between the two versions

The previous version of this page said the aircraft lined up and would not land. That is no longer true, and the table below is the whole argument for why a measured envelope is worth the trouble: it said exactly where to look, and the next version fixed most of it.

Identical 162 starts, scored identically
 previouscurrent
Reached the runway1847
On the aim point510
Touched down off the side of the runway371
Gave up (go-around)2317
Median error from the centreline16.6 m4.2 m
Flared, of those that touched down55 of 139118 of 145
Median sink at touchdown341 fpm195 fpm
Median touchdown past the threshold−157 m−145 m

Side to side: solved

Started anywhere from on the centreline to 400 m beside it, the aircraft now arrives within 2.7 to 5.7 m of the centreline — the starting error still has essentially no effect on where it ends up, but the error itself has come down by a factor of four. Touchdowns off the side of the runway went from 37 to 1.

Started offMedian error from the centreline at touchdown
centred2.7 m
50 m3.8 m
100 m5.2 m
150 m3.5 m
200 m4.6 m
300 m5.7 m
400 m5.7 m

The bias that was visible before is still there, and still systematic: 118 of 145 touchdowns landed left of the centreline, in the steady 8 kt crosswind from the right that every start was flown in. It is now a 4 m offset rather than a 17 m one. A systematic error is a far better problem to have than random scatter, because it points at a single correction rather than a tuning exercise — and shrinking it without removing it is what you would expect from a fix that improved control rather than addressed the cause.

Along the runway: it depends entirely on the length of the final

Of 145 touchdowns, 47 reached the paved surface, up from 18 on the same starts. The median still lands 166 m short of the aim point, with the middle half falling between 194 m short and 8 m past it. But that single median hides the finding, because the behaviour splits hard by how far out the approach began.

By distance from the runway at the start — current version
StartedMedian past the thresholdReached the runwayFlaredGave up
1.5 nm0 m21 of 5441 of 4113
2.0 nm−135 m17 of 5443 of 504
2.5 nm−163 m9 of 5434 of 540

From 1.5 nm it lands on the threshold — median 0 m, and every single touchdown flared. From 2.5 nm it lands 163 m short, which is where the previous version landed from everywhere. The longer the final, the shorter the landing, and the failure is no longer the flare: most of the short landings are flared. The aircraft is flying a path that is already low by the time there is anything to flare for, and the further out it starts the longer it has to get there. That is a glidepath problem on the long final, and it is the next thing being fixed.

undershoot terrain THR aim point median 40 L20 L020 R metres from centreline −300−200−100THR+100+200 metres past the threshold
Reached the runway — 47 Missed the runway — 98 On the aim point — 10
Every touchdown, all 145 of them. Plan view of the first 300 m of runway, drawn to the same scale along and across, so the shape of the scatter is the real shape. Both axes are measured from the runway: sideways from the centreline, lengthwise from the threshold, with the aim point the standard marking 21 m beyond it. The cloud is narrow and long, and it is narrower than it was: a Cessna 172 spans 11 m, and almost every point now sits within half a wingspan of the centreline, while the same points still spread over more than 200 m of runway and undershoot. The shape of the remaining problem is entirely in one axis. A human-written reference controller, flying the same navigation, reaches the runway in 5 approaches out of 6, so this spread is not a navigation limit.

The flare was the last finding. It is no longer the problem.

The previous version flared on 55 of 139 touchdowns. This one flares on 118 of 145 — and on every single one of the 41 touchdowns from 1.5 nm. Median sink across all touchdowns fell from 341 to 195 ft/min.

The gap between flaring and not is still stark, which is what tells you the flare itself is working when it fires: where the aircraft flares it arrives at a median 140 ft/min, and where it does not it hits at 452 ft/min, more than three times harder. Twelve touchdowns were hard enough to be recorded as such.

But the flare no longer explains the short landings, because most of the short landings are flared. It arrives low, flares correctly, and touches down on the grass before the threshold. That moves the open problem back up the approach, onto the vertical path on the long final — a different and more tractable question than the one this page carried a day ago.

GENunseen fields

It does not matter whether it has seen the airfield

One of the three runway ends was part of what the system learned from. The other two were deliberately kept out. If it were memorising airfields rather than flying approaches, that would show here. It does not — the trained runway is not better on any measure, and the held-out field that improved most between the two versions, Kuala Terengganu, went from 3 touchdowns on the runway to 17.

54 approaches at each runway end — current version
Runway end Learned from Reached the runway On the aim point Gave up Median miss
Seletar 03trained1724164 m
Kuantan 18held out1334166 m
Kuala Terengganu 04held out1759167 m

Three airfields is not a generalisation claim, and it is not offered as one. It is a negative result that matters: whatever is causing the short landings, unfamiliarity with the airfield is not it.

TYPEother aircraft

Put it in an aeroplane it has never flown

Until this week the only honest thing this page could say about other aircraft types was that nothing was known. Something is now known. The short version is that it does not transfer — but the reason is more useful than the headline, and a good part of it is not the system's fault.

Eleven types flew the same approach: centred, on the glidepath, from 2 nm, once at each of the three runway ends, in the same 8 kt crosswind. 33 approaches. Only the flight dynamics changed — what the aircraft sees out of the window is identical. The types run from a 50 kt Sopwith Camel to a 120 kt business jet. The system has only ever flown a 65 kt Cessna 172.

The caveat has to come first, because without it these numbers mislead. The inner loops that hold speed, glidepath, flare and go-around are tuned for a 65 kt Cessna 172, and were not retuned for any of these types. So this tests the whole stack on an unfamiliar aeroplane, not the vision and decision part on its own. To separate the two, the human-written reference controller — the same one used elsewhere on this page — flew the identical 33 approaches through the identical loops.

33 approaches each — same start, same three runways
Type Approach kt Reference This system What it did
Cessna 172 (i)652 of 30 of 3touched down 162 m short to 3 m short
Cessna 172 (ii)650 of 32 of 35 and 27 m past the threshold; one went around
Cessna 172 (iii)650 of 30 of 382 to 196 m short
Cessna 182702 of 31 of 3one landing 41 m past; two went around
Pilatus PC-7903 of 32 of 3two long, one 151 m short
T-6 Texan II900 of 30 of 3119, 129 and 131 m short
Piper J-3 Cub503 of 30 of 3landed long, 24–30 m off the side, all three
Sopwith Camel501 of 32 of 358 and 345 m past; one went around
Cessna T-371001 of 30 of 3went around every time
Cessna 310951 of 30 of 3went around, then lost control
Global 50001202 of 30 of 3went around, then lost control
All 33—157on the runway

“On the runway” means past the threshold and inside the 23 m half-width. Three different flight-dynamics models of the Cessna 172 were flown; they do not behave identically, which is itself worth knowing.

Half the ceiling is the rig, and half of what is left is real

The reference controller managed 15 of 33 through the same loops. That is the ceiling this rig imposes on an untuned type, and it is low. Against that ceiling the system managed 7 — about half as good as a controller that was written by hand and also never tuned for these aeroplanes.

Where the reference also fails, the system is not the limiting factor. Nothing landed the T-6: the reference put it 174 to 178 m short on every run, and the system actually did better, at 119 to 131 m short. Where the reference succeeds and the system does not, the finding is real — and there are two of those.

The Cub: a slow aeroplane in a crosswind

The reference landed the J-3 Cub on the runway all three times. The system put it 24 to 30 m left of the centreline on all three, just off the side, at all three airfields. At 50 kt the same 8 kt crosswind pushes an aeroplane much further than it does at 65 kt, and the corrections learned at 65 kt do not take it out. This is the clean version of the lateral bias reported further up this page: same direction, same cause, six times the size.

Every go-around gave a reason that was not the reason

13 of the 33 approaches went around, between 170 and 480 ft. Every one of them recorded the same cause: the runway was not in sight. In all 13, the runway had already been sighted on the approach.

So the decision to abandon is not being driven by whether the aircraft can see the runway. It is reacting to an approach that feels unfamiliar — a speed and a sink rate outside anything it has flown — and reporting the one reason it has available. The behaviour is arguably correct; giving up on an aeroplane you cannot fly is the right call. The explanation it gives for doing so is wrong, and that is the part that matters. It is the same failure as the navigation system that got lost and did not know it: the system's account of its own behaviour did not match what it actually had.

What the rig could not do

The Cessna 310 and the Global 5000 lost control after going around — a recovery loop built for a 65 kt single cannot fly a 95 kt twin or a 120 kt jet. Those rows are a rig limit, not a result. Of 39 types screened with the reference controller, only these 11 could be flown through the C172 loops at all; the rest landed far short, drifted off the side, floated, or could not be loaded. Per-type tuning comes before any of them mean anything.

Read together: the lateral behaviour transfers wherever the approach speed is near what it learned, and almost nothing survives a large change in speed — the Cub fails laterally for a reason that is really about speed. And the go-around trigger needs to be grounded in what the aircraft can actually see rather than in how unfamiliar the approach feels. None of that is visible from 162 approaches in one aeroplane, which is the argument for flying the ugly experiment.

BOXon the hardware

Running on a box that could fly

Everything above was produced with the decisions being made on a standalone embedded module — a Jetson, the class of hardware that would actually be bolted into an airframe — not on the desktop machine running the simulation. It answered in about half a second per decision, which is inside the budget an approach needs, and it made 34,427 of them across the campaign — a median of 228 per approach — over 6 hours and 8 minutes, 162 approaches and three airfields.

That matters more than the accuracy numbers do. A result that only exists on a workstation is a result about a workstation. This one has already survived the move to hardware that could be carried.

It was not an unattended run, and the earlier claim on this page that it was has been withdrawn. The simulator host needed two interventions: a launch that died seven minutes in and had to be restarted, and a window that opened over the simulator late in the run and was cleared in about twenty seconds. Both were the Windows machine driving the simulation. Neither was the decision box, which ran the whole campaign without a restart — but the distinction only counts for something if the failures are reported too.

WXweather limits

Where the weather stops it

Navigating by what the aircraft can see means the weather is not a nuisance variable, it is a hard boundary. These are the measured edges.

Holds

Daylight, up to 20 kt of wind with gusts, scattered cloud, 10–15 km visibility. Position held to 11–52 m at the start of the approach on both routes. This is the condition set the navigation figures come from. The approach matrix is a separate campaign, flown in a steady 8 kt crosswind from the right at every start.

Degrades — low sun

At sunset and at daybreak the aircraft could place itself less than half as often. It entered the cruise 650–700 m out of position and still found the runway, but only after going around once. The cause is mundane and fixable: the reference imagery was captured at midday, so the shadows do not match.

Hard limit — broken cloud

Broken cloud in the late afternoon defeated the navigation outright. With too little ground in view, the aircraft never established where it was, and the two flights ended 5.9 km and 45–49 km from where it believed it was. Both were stopped and excluded from every number on this page rather than quietly dropped. A 25 kt wind was present in the same runs, so cloud is not isolated as the sole cause.

The finding that matters most from those two flights is not that the aircraft got lost. It is that it did not know it was lost. Its own confidence estimate had degraded to kilometres and it kept flying the plan regardless. A navigation system that cannot declare its own loss of integrity is not a candidate for anything, and making it say so now sits ahead of accuracy work.

SCOREhow it is counted

How these numbers are counted

The figures above are worth what the counting behind them is worth.

  • Every start is reported. 162 planned, 162 flown, 162 scored, for each of the two versions. Nothing was re-flown for a better result and no approach was dropped.
  • The two versions are compared on identical starts. Same 162 cells, same crosswind, same scoring, same decision box. The older version was re-scored on this matrix rather than quoted from its own larger campaign, so no comparison on this page is between different sets of flights.
  • One attempt each. A go-around ends the approach and counts as a failure to recover, rather than being retried until it worked.
  • Scored against a published datum, not a judgement call: distance from the standard aim-point marking 21 m past the threshold, plus error from the centreline and sink rate at touchdown.
  • Errors applied one axis at a time, so an outcome can be attributed to a specific error rather than a combination.
  • Trends, not cells. Each combination was flown once per airfield, which supports the shape of the envelope but not a failure rate for any individual cell. Repeats at the boundary are outstanding.
  • Invalid runs declared, not deleted — see the two excluded weather flights above, and the two simulator-host interventions recorded under the hardware.
  • Trained and held-out airfields reported separately, never blended into one number.
SCOPEwhat this is not

What this is not

Flight status
Simulation only. Nothing here has flown on an aircraft. A commercial flight simulator over photographic satellite imagery of real terrain, with modelled wind and turbulence.
Not yet a camera
The embedded module receives the simulator's out-the-window view over a local link, not a real camera feed. Lens, exposure, vibration and motion blur are all still ahead.
Aircraft
Cessna 172. One airframe type, one weight, flaps and power handled by the aircraft's own systems. Every figure above except the airframe tour is a Cessna 172. It does not transfer to other types — 7 landings on the runway out of 33 across eleven aeroplanes, against 15 for a hand-written reference controller through the same untuned loops. Nothing here claims otherwise.
Position truth
Error is measured against simulator truth, not a surveyed ground reference.
Known optimism
The reference imagery and the imagery flown come from the same source, which makes recognising places easier than reality will.
Sample size
Three runway ends, one flight per cell. Enough to show the envelope's shape, to rule out airfield familiarity, and to compare two versions on identical starts; not enough to quote rates.
Provenance
Personal project, personal hardware, personal time. Public imagery and published airport data. No employer data, systems or material.

Method, architecture and implementation are deliberately not described on this page.