A model trained on simulated sonar imagery, or an analysis run on it, inherits every error the simulator makes. If the synthetic images carry a sidelobe structure, a level against range or a speckle spectrum that no real sonar would produce, a detector learns it and then fails on field data for reasons that are hard to trace. An acoustician asked to trust results built on simulated data needs the evidence that would be asked of a real sonar before it was accepted: measured resolution, sidelobes and spectra that agree with what the physics of the instrument predicts.
This page is that evidence for ApertureLab's point-scatterer simulator and its time-domain back-projection beamformer. It images a deliberately simple sonar, one element on a straight track, and then a realistic 36-channel array in continuous motion, and it compares every measurement with a prediction computed from the instrument's parameters. It also compares the statistics of simulated seafloors with the scattering models used in the field. The contents list below leads to each section, so the page can be read from top to bottom or entered at the one result a reader needs. The same measurements run as a regression test, so a change to the code that moves any of them fails before it is released.
Contents
The instrument
Every feature a working sonar adds, such as a multi-element array, platform motion while the echo is in the water, gain control or glint suppression, adds a term to the prediction and a place for an error to hide. The reference instrument therefore has none of them. One element transmits and receives, the platform stands still for each ping, and the beamformer integrates a fixed angular sector under a Hann window. What remains is the physics of a chirp, a beam and a synthetic aperture, whose predictions can be computed without the simulator. The array in continuous motion is checked later against the same predictions.
Point targets
Six deterministic point scatterers sit on a silent floor at 20 to 95 m ground range, staggered along-track so that their sidelobes do not overlap. Each is beamformed on its own 4 m patch at 6.25 mm pixels, sinc-interpolated four times, and measured for the position of its peak, its −3 dB width on both axes, its peak sidelobe ratio and its integrated sidelobe ratio. The peak sidelobe ratio (PSLR) is the level of the highest sidelobe relative to the peak. The integrated sidelobe ratio (ISLR) is the energy outside the mainlobe relative to the energy inside it, taken over the 2 m square around the target with the mainlobe bounded by the first nulls of the two principal cuts.
The predictions are computed for the weighting this beamformer applies, not read from a textbook formula. Along-track, the aperture is shaped by the Hann window on the integrated sector and by the element's own pattern, applied twice because one element transmits and receives. The predicted response is the Fourier transform of that weighting over the along-track wavenumber 2k sin θ, summed across the chirp band with the exact range to every ping. In range, the prediction is the compressed pulse of the transmitted chirp after the beamformer's own front end: a Hann-windowed matched filter, and a low-pass filter on both the data and the replica. The closed forms are drawn as references: d / 2 = 15.0 mm along-track, and 1.44 c / 2B = 18.0 mm of slant range against 18.07 mm from the numeric model, with a Hann window alone giving a first sidelobe of −31.5 dB.
Every peak lands under 0.05 mm from the true position. The along-track width narrows from near range to far, and the prediction follows it, because the beamformer gates the aperture on the angle in the horizontal plane while the along-track wavenumber a ping contributes is set by the three-dimensional look direction. At 20 m, where the platform looks down at 27°, the same sector spans 10 percent less wavenumber than at 95 m, and the focused point is wider in proportion. The regression test allows 0.95 to 1.05 times the prediction; the largest departure measured on either axis is 0.3 percent.
The faint bow-tie around each target in the overview is the integration pedestal: the target's own echo, summed with the wrong phase into every pixel that shares its aperture, so it spans the synthetic aperture length at that range. Measured half a metre to 0.9 m from each peak along its range line, it lies 55 to 57 dB below the peak. It shows here because the floor is silent and the display spans the whole dynamic range; whether it shows over a real seafloor depends on how bright the target is against the bottom return.
Peak level against range is reported rather than gated,
and it has a prediction of its own. With gain control off, the simulator applies
spherical spreading both ways, seawater absorption of 0.067 dB/m at
300 kHz (Thorp's formula, as simulator/attenuation.py implements it), the
sin(grazing) factor a point on a Lambertian floor carries and the element's elevation
pattern; the beamformer adds the coherent gain of an aperture that lengthens with range.
The model evaluates all of them at the six targets and fits one slope, set out term by
term in the figure.
Flat sand
A uniform sandy bottom is speckle: its image carries no structure of its own, so its two-dimensional Fourier transform is a picture of the imaging system. In wavenumber space the support is an annular wedge. The radius is the two-way wavenumber 2k, so the radial band is set by the chirp, 270 to 330 kHz, and the angular extent is the range of aspect angles the beamformer integrated, ± 6.35° about broadside. Because the image is in ground range, the wedge is foreshortened by the cosine of the depression angle, which is why it is measured on 6 m patches at six ranges rather than on the whole image.
Inside the wedge the energy is not flat. Across angle it is shaped by the Hann window the beamformer puts on the sector and by the element's pattern, applied twice because one element transmits and receives; across radius it is shaped by the Hann window on the matched filter. Those two curves are drawn in white over the measurements, and the figures below read both at −3 and −10 dB. The measured −10 dB half-angle is 3.12 to 3.31 degrees across the six patches against 3.25 predicted, and the whole angle curve fits the prediction to 0.53 dB rms or better above −20 dB.
Back-projection rotates every sample by the carrier phase of the pixel's own delay, which leaves each focused pixel carrying a phase ramp along range, one turn every 2.5 mm of slant range. The beamformer removes that ramp as its last step by default, so its output is at baseband and every patch on this page is measured as it comes out, with nothing subtracted afterwards. The figure shows the check on a strip across the whole swath: with the ramp kept, the centre of each block's range spectrum follows the predicted fold of the carrier to 6.7 rad/m rms, and with it removed every block sits within 3.4 rad/m of zero. The Method appendix explains the convention and why it matters for anything that filters or resamples the complex image.
Seafloor statistics
The statistics of a sonar image matter as much as its resolution: the tail of the intensity distribution sets the false-alarm rate of any detector run on it. The field describes that tail with the K-distribution, whose shape parameter α is tied to the number of scatterers in a resolution cell (Abraham and Lyons 2002) and to the modulation of scattering strength by seafloor slope (Lyons, Olson and Hansen 2022). A simulator used to train or test such detectors should reproduce these statistics from its physics. Because the simulator knows every input a survey has to infer, each statistic below is compared with a prediction made from those inputs, not fitted to the data.
The measure is the scintillation index, SI = 〈I2〉 / 〈I〉2 − 1, the normalised variance of the intensity I. It is 1 for Rayleigh speckle and 1 + 2/α for a K-distributed intensity. The first test images flat sand with the single-element instrument above, 15 to 25 m along-track and 20 to 45 m in ground range, at seven scatterer densities from 200 to 50,000 per m2. The simulator places scatterers at random with complex Gaussian amplitudes, and for such a field SI = 1 + 2/(ρ·Aeff) exactly, where ρ is the density and Aeff = (∫|h|2)2 / ∫|h|4 is the effective area of the point-spread function h. Aeff is measured from six point targets at the same ranges on the same track, 7.46 to 8.69 × 10−4 m2, so the equivalent shape parameter α = ρ·Aeff is known for every image.
At all seven densities the prediction lies inside the measured 95% block-bootstrap interval, and the measured SI is within 2% of SI − 1. The whole distribution agrees as well: the probability that a pixel exceeds a threshold follows the exact point-scatterer model down to one pixel in a million. From about one scatterer per resolution cell upward that distribution cannot be told from the K-distribution with the same α. Below it the K curve overstates the far tail, because the K-distribution describes clustered scatterers and these are placed independently.
Lyons, Olson and Hansen measured α = 2.3, 4.0 and 7.9 for the speckle at three sites with three different systems. For this instrument those values correspond to about 2,900, 5,100 and 10,100 scatterers per m2. The density is a numerical parameter, so this shows that the simulator can produce the speckle statistics measured at sea, not how a particular seafloor produces them. On a real seafloor heavier tails come from its relief, its patchiness and its ripples. The measurements that follow take those up in turn: the scintillation index against range on randomly rough sand, compared with Eq. 6 of Lyons, Olson and Hansen (2022), then resolution, ripples and boulder fields.
Johnson (2009, Ch. 4) found that the statistics of real SAS images depend strongly on resolution, by reprocessing the same data at coarser resolution through sub-banding of the image spectrum. The point-scatterer result predicts that dependence exactly: a coarser resolution cell holds more scatterers, so SI falls toward 1 as 1 + 2/(ρ·Aeff). The 2,000 per m2 image was sub-banded to one half, one quarter and one eighth of its bandwidth on both axes, and the same filter was applied to the six point-target responses to give each Aeff. The measured SI falls from 2.27 to 1.02 as Aeff grows fiftyfold, and every point matches its prediction to within its 95% interval.
The second test puts relief under the same instrument: power-law sand with spectral exponent 3.25, the value Lyons, Olson and Hansen (2022) use at all three of their sites, imaged 20 m along-track from 5 to 100 m in range at 10,000 scatterers per m2. Local slope changes the grazing angle of each patch of seafloor and so its scattering strength, and that modulation raises SI above the speckle value, most at low grazing angles where the scattering law is steepest. Their Eq. 6 predicts SI against range from the rms slope and the scattering law. Here the law is the simulator's own Lambert law, the rms slope is measured from the generated heightfield rather than fitted, and the speckle SI comes from the first test, so the prediction has no free parameter. The model leaves out shadowing and multiple scattering, which the paper notes matter once the grazing angle nears the rms slope, so the comparison stops where the grazing angle falls to twice the rms slope.
Beyond the range where the model holds, the simulated SI departs from Eq. 6 by up to 10% from one range band to the next, and the departures have a physical cause the simulator reproduces. A power-law surface carries large undulations that tilt whole patches of seafloor toward or away from the sonar, which moves the local mean grazing angle; Lyons, Olson and Hansen show the same effect in their Tellaro data, where it produces large fluctuations in SI with range. Eq. 6 assumes a flat mean seafloor. Evaluating it with the tilt measured under each range line brings the rougher of two surfaces with the same rms height (6.1° of rms slope) from a 6.5% rms error to about 2% over 20 to 75 m. At 90 to 100 m on that surface the grazing angle equals the rms slope, and the measured SI exceeds even the tilted prediction: facets hidden behind the slopes in front of them return nothing, which is the shadowing the paper expects in that regime and leaves out of the model.
Intensity is normalised by a smooth fit of its mean against range before the statistics are taken. Normalising each range line by its own mean biases SI toward 1, as Johnson (2009, Sec. 2.3.2) cautions; a first analysis did so and ran 1 to 13% low in SI − 1.
A realistic instrument
A single element in stop-and-hop is the easiest case to predict and the least like a working sonar. The second instrument is a 36-channel array of 30.0 mm elements on a vehicle moving continuously, over the same track, the same seafloor and the same six targets, with the beamformer integrating the same ± 6.35° sector under the same window and compensating the motion within each ping. The width predictions change only through the lattice of phase centres, because the element and the chirp are the same. Every summary figure in the Point targets and Flat sand sections draws both instruments, the single element as circles and the array as hollow squares, so the two can be compared at every range. Every array peak lands under 0.05 mm from the true position. Its widths depart from their predictions by 0.3 percent at most, against 0.3 percent for the single element. The carrier walk measures 7.7 rad/m rms from its predicted fold with the carrier kept and 3.3 rad/m from zero with it removed, against 6.7 and 3.4 for the single element. The two instruments differ in these settings:
At 20 m the array's highest along-track sidelobe is −43.9 dB against −43.6 dB for the single element. The vehicle advances 225.0 mm per ping, 15.00 phase-centre spacings, and the beamformer keeps 15 channels of each ping, 15 phase centres. Consecutive pings tile the aperture exactly. A phase centre summed in two pings, or a gap between them, would weight the aperture with the ping period and put a grating lobe at λR / 2a along-track; with the pings tiling the aperture there is none. The same numeric model as for the single element, given the array's phase-centre lattice, predicts the sidelobe level at every range, and the figure compares it with the measurement.
The flat-sand angle curve fits its prediction to 0.90 dB rms at 20 m for the array against 0.53 dB for the single element. The beamformer evaluates its Hann window at the phase centre of each transmit and receive pair, the midpoint of the transmitter at transmit time and the receiver when the echo arrives. An earlier version evaluated it from the receiver; with the transmitter at the array centre, the kept receivers all sit 105 to 525 mm ahead of the transmitter, and in continuous motion each has moved a further v · 2R / c while the echo was in the water, which put the window a fraction of a degree forward of broadside and the misfit at 1.9 dB at 20 m. The array's misfit is still larger than the single element's at every range, by about 0.3 dB, and the model, drawn dashed, accounts for only part of it. The remainder is not yet explained and is listed under what this page does not show. The model's own residual for the single element at 20 m comes from the prediction, which evaluates the element pattern in the horizontal plane rather than along the three-dimensional look direction.
Redundant phase centres
Micronavigation for synthetic aperture sonar rests on one geometric fact. With the transmitter at the centre of a multi-element array, each transmit-receive pair behaves like a single element at their midpoint, the phase centre. If the vehicle advances a whole number of phase-centre spacings between pings, the next ping re-occupies some of the previous ping's phase centres, and the echoes recorded there should be the same up to whatever the vehicle did in between. This is a constraint on the sonar, the ping rate and the speed together. The array above at 1.5 m/s advances 15.00 spacings per ping, so this check runs it at 1.53 m/s and 3.0 Hz, exactly 34 spacings, and consecutive pings share two phase centres.
Each channel is low-passed and matched-filtered as the beamformer does it, the record is cut into blocks, and each pair of blocks goes through the beamformer's own micronavigation estimator: a coarse lag from the correlation magnitude, refined by the carrier phase at its peak. The refined delay is reported as it comes, with nothing subtracted.
The delay should show only geometry and motion. A shared phase centre is formed by a transmitter and receiver L apart, whose two-way path is L²/4r longer than the phase centre's monostatic one, and because the vehicle keeps moving while the echo is in the water the separation that counts is L + vt. The predicted delay of one ping behind the other is the difference of the two receivers' excess paths: a 1/r curve for each pair and a motion term that is the same at every range. Summing the shared channels of a ping needs one step first, because that excess differs between channels by several radians of carrier phase at near range, so each channel is advanced to the subarray's centre channel before the sum. The measured delay follows its prediction to a few nanoseconds, and every window of every ping pair correlates well above the level the micronavigation solver accepts, which is the condition for estimating sway and heave from the data.
Each shared pair alone correlates at a median of 0.997 (minimum 0.996) and follows its predicted delay to 2 ns rms, 7 ns at most. Aligned and summed, the shared channels correlate at a median of 0.997 against 0.992 summed as recorded, and their delay averages −0.583 µs against −0.584 µs predicted, 2 ns rms. Over the whole track, 194 ping pairs by 88 windows correlate at a median of 0.997, 100.0 percent of them above the 0.66 acceptance, and the delay follows the prediction to 3 ns rms.
The check is set up as follows.
Scope
The checks on this page establish that the point-scatterer simulator and the back-projection beamformer agree with the physics of the instrument they describe: resolution, sidelobes, level against range, spectral support and the delays between redundant phase centres all land on predictions computed without the simulator. They do not establish everything a user of the imagery might assume, and the limits are these.
Method and gates
benchmarks/scenes/,
validation_flat_sand.yaml and validation_psf_points.yaml,
with identical sonar, motion and beamformer blocks; a test refuses to let them drift.
The array has its own pair, and a third scene, validation_muscle_rpc.yaml,
for the redundant-phase-centre check.psf_metrics.islr_2d_db, tested against the
closed forms for uniform and Hann apertures). Results published before
2026-09-22 carried a figure labelled ISLR that was the brightest single pixel outside a
1.5-width box, which equals the range PSLR; the page omits the column until a
measurement run has recorded the true figure.simulator/validation.py. Range:
range_psf_prediction builds the transmitted chirp and passes it through the
beamformer's front end, a zero-phase low-pass designed with a Hamming window at half the
bandwidth on the echo, and the Hann-windowed matched filter low-passed the same way; the
compressed pulse gives 18.07 mm at −3 dB in slant range and a first
sidelobe of −31.5 dB, projected onto the ground by R / y.
Along-track: along_track_psf_prediction sums, over every phase centre that saw the
target and over the chirp band, the echo the simulator records (two-way element pattern
along-track and in elevation, 1 / R², absorption 0.0669 dB/m by Thorp's
formula, the sin(grazing) factor) rotated back the way the beamformer does it, with its
Hann gate evaluated at the phase centre, as the beamformer does. The same sum at the peak gives the level
prediction. scripts/validate_sanity.py --theory-only recomputes every
prediction from the scene files without simulating.Back-projection focuses a pixel by sampling every ping at the pixel's own delay and rotating the sample by +2πfcτ. On a scatterer the rotation cancels the echo phase and the pings add coherently; one pixel over in range, a residual phase of 2kΔr is left behind. A focused point is therefore the compressed pulse times a carrier whose phase advances one turn every 2.5 mm of slant range, and the range spectrum of the image sits at 2k cos θdep, about 2248 to 2501 rad/m across the swath, rather than at zero. Each pixel's phase is then referenced to its own position, which suits interferometry and change detection and leaves every magnitude product unchanged.
The carrier matters to anything that treats the complex image as band-limited about zero along range: a spectrum, a sinc interpolation, a complex filter or a resampling. Sampled at 6.25 mm, the carrier folds into the ±503 rad/m pixel band, and because cos θdep changes across the swath, the folded position moves with range. The beamformer's last stage removes it: for every pixel it takes the minimum two-way delay from the platform track to that pixel, with heave, sway and curvature included, and multiplies the finished pixel by exp(−j 2πfcτref). The reference depends on the pixel and not on the ping, so it commutes with the sum over the aperture and is applied once after all pings are in, and the interferometric phase between two images on the same grid is unchanged. The removal is on by default, and every SLC records whether it was applied.
The figure shows the whole overview cut into 11 by 15 blocks of 6.4 m, each block's two-dimensional spectrum drawn in its own place, with the carrier kept on the left and removed on the right. At 25.0 mm the band is wider than the pixel window and folds onto itself, which is why the measurement in the Flat sand section is made on a strip beamformed at 6.25 mm.
References