21/09/2026

Which Bulletins Are Worth Reading

Ilma is my self-built Navtex decoder for the Raspberry Pi. After powering its antenna properly and a run of decoder fixes, it began hearing stations it had never heard from my home berth: Oostende, about 140 nm away, and Pinneberg, about 160.1

Most of what those two deliver is not readable. The decoder marks every character it could not recover with an asterisk, per the ITU recommendation,2 and a weak Pinneberg bulletin arrives looking like this:

SA21
$,::-HAMBURG
241110 UTC AUG *6
NDV.*WARN. NO. 481
*ERPAN  GHT. WEST OF OFFHORE WINDPARK 'SANDBANK'.
NOT***ENTIYIT* **STRUCTION IU*55-21,62N 00*-31,71E, WATERDEP*T 37M AND
 5-19,43N 006-30,,3E, WATERDEPHT 37M
DANGER FOR KNSHERY.

Which is, on inspection, useful: an unidentified obstruction near the Sandbank wind farm, at a stated position, in 37 m of water. Many of its neighbours are nothing at all. Can the receiver tell the difference by itself?

Three candidates

The erasure rate of the body is free: the decoder already counts the characters it gave up on, for all 1,099 messages in the archive. The per-message signal-to-noise ratio is newer, measured over each message’s own time on air, and only attached to rows decoded since 20 September.3 The third was the interesting one: Ilma reads the boat’s position off the NMEA 2000 bus, so it could compute the distance to each transmitter and treat far-and-weak as suspect.

Comparing them needs an independent verdict on which bulletins are genuinely unreadable. Scoring each message against a vocabulary built from the clean captures failed usefully: it called a good Den Helder warning about a discontinued light buoy 39 per cent readable, because coordinates and buoy names are not vocabulary. Scoring letter sequences works, using a character trigram model fitted to the 350 clean bodies in the archive.4 YELLOW LIGHTBUOY scores well whatever the word; ZBEDARNING,1$ does not.

Receiver operating characteristic for three candidate tests.

Figure 1. Receiver operating characteristic (ROC) for three candidate tests. Each curve is one candidate test, traced across every threshold it could take. Up is unreadable bulletins correctly hidden; right is readable ones hidden by mistake.
The erasure rate scores 0.931 over the whole archive and 0.974 over the smaller set that also carries a signal-to-noise measurement. Tone signal-to-noise scores 0.941 on that same set and is the weaker test of the two, despite being the more expensive one to obtain.
The open circles are each curve's statistical optimum. The filled one is where the filter actually runs.

The erasure rate wins, applies to every message ever recorded, and needs no new plumbing.

Why distance would have added nothing

Signal-to-noise looks respectable in Figure 1. The measurements underneath show what it is really doing.

Per-message tone signal-to-noise against body erasure rate, by station.

Figure 2. Every message carrying a per-message signal-to-noise measurement, against the erasure rate of its own body. Filled markers read as intelligible text, hollow ones do not.
Den Helder occupies a tight cluster at 22.1 to 23.0 dB, and many of its messages sit on top of one another at zero erasures. Everything else falls at or below 10.4 dB. Nothing at all lands in the 11.2 dB between them.
The best signal-to-noise threshold the optimisation could find, 21.66 dB, falls inside that empty band. It is not measuring quality. It is asking whether the message came from Den Helder.

That empty band answers the distance question. The threshold that scores well is a station-identity test wearing a quality test’s clothes, and distance sits a step further from the measurement again, because distance is what causes the signal-to-noise difference. It would have given the filter a noisier copy of what it already had, by a longer route. The position is also mostly absent: it reaches Ilma from a separate process reading the Raymarine bus, and 154 of 1,010 broadcast windows carry one. On the day I ran this, none of the 45 did.

Navtex transmitter stations in Northern Europe.

Figure 3. Navtex transmitter stations in Northern Europe with their nominal range in nautical miles. Station P is Den Helder, S is Pinneberg.

Choosing the threshold is not choosing the test

The obvious cut is the point that optimises both properties at once, about 5 per cent, and setting it there would have been a mistake. That optimum treats a hidden warning and a displayed scrap as equally bad. On a boat they are not: one is an annoyance, the other a navigational warning I do not get. Against the 653 bulletins the default view shows, a 5 per cent cut hides 114 messages and 37 of them are legible.5 Among the casualties, at 5.2 per cent erasures:

JA59
230925 UTC *UL 26
KAL *AV WARN 170/26
SOUTHEASTE*N BALTIC
SHI*S EXER*ISES*3121*0 UTC JUL THRU 312100 UTC AUG*IN AREA TEMPORARILY DANGEROUS TO SHIPPING BR-42

At 15 per cent the filter hides 37 and loses none, while still removing 37 of the 94 unreadable ones. That is the cut I shipped: the least sensitive setting that costs nothing, which is a different objective from the one the optimisation was solving.

One correction followed. Measuring across the whole stored message hid a Pinneberg warning about ammunition found north of Norderney: 247 characters at 0.4 per cent erasures, with 632 characters of noise attached to the back, giving 28.7 per cent overall. The message had never ended. A Navtex bulletin closes with NNNN, and here the four N’s arrived as N, a line feed, and one more N. The decoder tolerates one wrong bit in a marker character; this one was four bits out, so the framer kept filling with whatever came next on the air. The filter now reads only the first 200 characters, and six bulletins came back, every one from Pinneberg or Oostende.

The failure no filter sees

Three messages arrived at 20 dB or better and are unreadable anyway. Two carry almost no erasures. The decoder was not unsure; it was confident and wrong, which is the one state a confidence measure cannot report. One is a search-and-rescue bulletin whose body is perfectly legible:

0-5Q
QIPTEE UTC SEP 26
MSI 239/26
WESTEREMS APPROACH
CAPSIZED SMALL SAILING YACHT
REPORTED ADRIFT
LENGTH 6 MTR
1 PERSON MISSING
LAST KNOWN POSITION 53-38.5N 006-11.4E
AT 18SEP1630 UTC
KEEP A SHARP LOOK OUT

The first two lines should carry the station identifier and a timestamp. Because they do not, the message has no station, serial or subject, so it files as unframed and sorts with the rubbish. Nothing about it is weak: 22.8 dB, zero erasures.

The cause was a bug I had carried for months. The demodulator slices each symbol into ten and tracks which slice sits at the centre. That index is a position on a ring, where slice nine and slice zero are a tenth of a symbol apart, and treating it as a plain index meant a wobble across the seam quietly inserted or dropped a bit. Navtex uses a code in which every valid character carries exactly four marks out of seven, so a bit-rotated character still carries four, passes the check, and is preferred by the error correction over the intact copy. The decoder emits a real letter that is the wrong letter, at any signal strength. It cost a Den Helder gale warning at 22 dB, which arrived as SEE ****P*ZBEDARNING,1$ while the bulletin four minutes later in the same recording decoded perfectly.

Which is the argument for fixing things in the decoder rather than detecting them afterwards. A filter tight enough to catch a confidently wrong bulletin would have to be tight enough to hide the weak, real, half-readable Pinneberg traffic that the whole summer’s work was about receiving. A receiver operating characteristic (ROC) is a good way to choose between tests. It cannot tell you that the thing you most need to catch is not in the data you are testing on.6

Footnotes

  1. Oostende is station T on 518 kHz, Pinneberg is station S. Distances are straight-line from my current position to approximate transmitter positions, rather than along the path the signal crosses, and are given to two figures because that is all they are worth: Ilma holds no transmitter coordinates, which is one of the things the third section is about. Den Helder, by contrast, is some 25 nm away, which is why its signal arrives at 22 dB and never varies. Pinneberg is the awkward one: it sits about 59 nm up the Elbe from the river mouth, so its distance and its signal strength have never been in a sensible relationship. ↩

  2. ITU-R Recommendation M.476-5 §3.2.5 provides that a receiver shall display an asterisk in place of any character the error-correction mechanism could not recover. Without it a bad body silently shrinks and the operator has no cue that anything was lost, which also makes every downstream measurement of decode quality meaningless. ↩

  3. Tone power against the two guard bands either side, measured over the message’s own span on air rather than over the whole ten-minute window. The window-wide figure was the earlier measurement and is misleading for short bulletins, which occupy a small fraction of it. Only messages decoded since 20 September carry the per-message value, which is why two of the three curves in Figure 1 are drawn on 60 messages rather than 1,099. ↩

  4. Add-one smoothed character trigrams over letters and spaces, fitted to the 350 bodies in the archive that carry a station identifier and no erasure markers, and scored as mean log probability per character, then scaled down by the fraction of the body the decoder erased. A message scoring below 0.60 counts as unreadable. The model is fitted on the same corpus it judges, which would be circular if the question were how well Navtex English can be modelled; here it only has to separate real text from bit-rotated text, and it is not the quantity any of the three candidate tests measures. ↩

  5. The 653 are the messages the default view returns: framed, carrying a station identifier, within the default subject and area filters. Unframed captures are already hidden behind a separate toggle, so they are not the population this filter acts on. Of the 653, the trigram score calls 94 unreadable. At 5 per cent the filter hides 114 messages and 37 of them are readable; at 10 per cent, 54 and one; at 15 per cent, 37 and none. ↩

  6. 18th-century Scottish philosopher David Hume pointed out that no amount of past observation can logically guarantee a future outcome. Immanuel Kant adds: Your data is already pre-filtered by your mind’s architecture; you are only ever testing your own assumptions, not reality itself. Thomas Bayes doesn’t predict what the unknown thing is – he ensures the model’s math never assumes the dataset contains the entire universe. In the 20th century, philosopher of science Karl Popper built on Hume’s insight to explain why data-testing cannot prove a theory true. He noted that empirical data can only falsify a claim, never permanently confirm it. Popper famously used the Black Swan problem: Millions of white swans in your data cannot prove the rule that all swans are white. A single black swan – which is absent from a historical dataset – instantly shatters the model. In contemporary philosophy and risk analysis, Nassim Nicholas Taleb applies this directly to modern data science, finance, and AI. Taleb highlights the danger of empiricism without domain awareness: relying strictly on backtesting, statistical models, or training data creates a false sense of security because data can only contain what has already occurred and been recorded. ↩