← ApertureLab

Why synthetic aperture?

A sonar image is sharp in two directions at once, and the two are limited by entirely different things. Sharpness down-range comes from the bandwidth of the transmitted pulse, and it holds up at any distance. Sharpness along the direction of travel is set by geometry: the size of the transducer against the wavelength it works at.

That second limit is the unforgiving one. It degrades in proportion to how far away you are looking, and no amount of transmitter power or receiver quality moves it. Defeating it is what synthetic aperture is for.

This page walks the idea from a single transducer to a synthetic aperture, one step at a time. The numbers throughout come from the array ApertureLab simulates: elements 33 mm long, driven at 300 kHz, where sound in water has a wavelength of 5 mm.

None of the argument is specific to sound. Radar arrived at it first, and synthetic aperture radar (SAR) is the same construction carried out with electromagnetic waves; every step below holds for either. This page works in sound, where the technique is called synthetic aperture sonar (SAS), because that is what the numbers describe. The last section sets out where the two modalities genuinely part company.

Step 1 of 7

Start with one element.

An element is a single transducer: the smallest piece of a sonar that turns an electrical signal into sound, and turns a returning echo back into a voltage. A microphone and a loudspeaker in one part, tuned to a narrow band and coupled to water instead of air.

It has a physical size. The element used throughout this page is 33 mm long. The sonar drives it at 300 kHz, and sound travels at roughly 1500 m/s in seawater, so one wavelength is 5 mm.

Those two lengths, the size of the element and the wavelength, decide almost everything that follows. Every array on this page, real or synthetic, is built out of elements like this one.

D 33 mm electrical drive one element acoustic wave, λ = 5 mm 300 kHz in seawater
One transducer element. Its length D and the acoustic wavelength λ are the only two numbers the rest of this page needs.

Step 2 of 7

One element cannot tell two things apart.

Sound leaving an element does not travel as a pencil. It spreads into a cone, and the angular width of that cone is set by the ratio of wavelength to element size:

θ ≈ λ / D = 5 mm / 33 mm = 0.15 rad = 8.6°

Now put a target in the water. The echo comes back at a time that pins down its range precisely, and a wideband pulse separates two objects a few centimeters apart in range no matter how far away they are. Range resolution does not degrade with distance.

Direction is another matter. All the echo tells you is that something is somewhere inside the cone. Two objects at the same range, anywhere across that cone, send back echoes that arrive at the same instant and add together. They are one measurement, not two.

And the cone widens with distance, so the ambiguity grows in proportion to range:

0.30 m
across the beam at 2 m
3.8 m
at 25 m
7.5 m
at 50 m
15 m
at 100 m
θ ≈ λ/D = 8.6° 3.8 m 7.5 m 15 m 25 m 50 m 100 m two targets, same range, both inside the beam; the echoes arrive together and add along-track range
Drawn to scale in both axes: the wedge really is 8.6 degrees wide, and the footprint really does reach 15 m across at 100 m. Range resolution stays fixed while along-track resolution degrades linearly with distance.

So the picture a single element produces is sharp in one direction and smeared in the other, and the smearing gets worse the further out you look. That asymmetry is the entire reason synthetic aperture exists.

A note on conventions

Beamwidth can be quoted as the half-power width, the first-null width, or the nominal λ/D. They differ by factors near one. This page uses λ/D throughout so the numbers stay comparable from step to step.

Step 3 of 7

Directionality is not a property of a transducer.

Take a second element and place it a distance d from the first, then add the two received signals together.

For a wave arriving straight ahead, the crests reach both elements at the same moment. The two voltages are in step, and the sum is twice either one.

For a wave arriving off to the side, one element is further along the incoming wavefront than the other. The extra distance the wave has to cover is d sin θ. When that extra distance is half a wavelength, the two voltages are exactly opposite and the sum is zero. The pair is blind in that direction.

Sweep the angle below and watch it happen.

geometry broadside θ A B d = λ amber: the extra distance the far element has to cover what each element receives A B A + B time → amplitude of the sum, against arrival angle 0 sum -90° -60° -30° 30° 60° 90°
24°
arrival angle
0.41 λ
extra path d sin θ
146°
phase difference
0.59 ×
sum, vs. one element

Drag the angle and watch B slide against A. Where the extra path reaches half a wavelength the two are opposite and the sum collapses to nothing; the curve underneath is that cancellation plotted across every angle, which is the beam pattern of a pair of elements. Spacing here is one wavelength, which puts the first null at a visible 30 degrees. The second peak at 90 degrees is a real ambiguity, and suppressing it is one of the jobs the extra elements in step 4 do.

Nothing about either element changes as you sweep. Neither one knows which way it is pointing, and neither one is more sensitive in any direction than it was before. The whole response comes from the phase relationship between two separated measurements.

This is the pivot the whole subject turns on. Directionality is something you construct out of separated measurements. It is not something a transducer has.

Step 4 of 7

More elements make one longer aperture.

Add more elements along the same line. More cancellation directions appear, and the one direction where everything still adds gets narrower. For N elements spaced d apart, the beam narrows to roughly λ/L, where L = N·d is the total length of the array.

The count is not what matters. The length is. Two arrays of the same overall length behave much the same whether they are built from 16 elements or 64; the extra elements buy freedom from ambiguity and signal-to-noise, not sharpness. To halve the beam you have to double the length.

94 cm L 50 m
8
elements
0.27 m
aperture length L
1.07°
beamwidth λ/L
0.94 m
footprint at 50 m

The beam and the 50 m range axis are drawn to the same scale, so the footprint you see is the footprint you get. The array itself is drawn about 14 times larger than scale, otherwise it would be a few pixels tall.

The real receive array in these simulations is 36 elements at 33.33 mm, so 1.2 m end to end. That takes the beam from 8.6 degrees down to a quarter of a degree, and the footprint at 50 m from 7.5 m down to 21 cm. Which is a large improvement, and still not enough.

Step 5 of 7

Then the physical array runs out.

Suppose you want along-track resolution roughly the size of the element itself, say 17 mm, out at 100 m. Rearranging δ = Rλ/L gives the array length that would take:

L = Rλ / δ = 100 m × 0.005 m / 0.017 m ≈ 30 m

Thirty meters of array. Rigid, held straight to a small fraction of 5 mm, towed behind a vehicle two meters long, through moving water.

2 m vehicle 30 m of rigid array, straight to a fraction of a millimeter to resolve 17 mm at 100 m
Vehicle and required array, drawn to the same scale.

This is not an engineering problem waiting on better materials. It is where real apertures end.

Step 6 of 7

Take it to the limit: build the array out of time.

Look again at step 3. Two elements a distance d apart were two measurements of the same wavefield from two separated points. Nothing in that argument required the two measurements to happen at the same instant, or to come from two different pieces of hardware.

The vehicle is already moving in a straight line, along exactly the direction you want the array to run. One element at position 1 during ping 1, and that same element at position 2 during ping 2, are two measurements from two separated points. That is all a two-element array ever was.

So do not build the array. Fly it. Fire a ping, move, fire again, and keep every complex echo. Afterwards, in software, sum the returns from all the positions with the phase each one carried, exactly as an array sums across its elements. The aperture is synthesized in time rather than assembled in metal, and it can be as long as you are willing to travel.

direction of travel synthetic aperture Rλ/D of travel one point on the seafloor ping 1 ping n each ping contributes one element; the summation happens afterwards, in software, with the phase of every ping preserved
Along-track spacing is exaggerated here for legibility. In the real geometry the whole run subtends the same 8.6 degrees seen in step 2.

The length available is not arbitrary. A point at range R stays inside the element's 8.6 degree cone for Rλ/D of travel: 7.5 m of aperture at 50 m range, 15 m at 100 m. The aperture the geometry hands you grows in exact proportion to range, which is the same proportionality that was ruining the resolution in step 2.

Then why does the real sonar still carry 36 receivers?

For sampling rate, not for sharpness. The synthetic array has to be sampled about every D/2 of travel or it aliases, and a single element would force a ping rate that caps how fast the vehicle can survey. Recording 36 receivers per ping lays down 36 samples of the synthetic aperture at once, so the vehicle can move faster between pings. The extra elements buy coverage rate.

Step 7 of 7

What you get, and what it costs.

Put the two results together. The aperture available at range R is Lₓₖₙ = Rλ/D. The resolution of an aperture that long is δ = λR / (2·Lₓₖₙ), where the factor of two comes from the round trip: the phase of a synthetic array turns twice as fast as a real one, because the path is out and back. Substituting:

δ = λR / (2 · Rλ/D) = D / 2 = 17 mm the range cancels

Along-track resolution becomes half the length of one element, at 10 m and at 100 m alike, and it does not depend on wavelength either. A shorter element gives a finer image, which is the reverse of what every optical instrument teaches. The same round-trip factor is why 15 m of travel replaces the 30 m of towed array from step 5.

1 cm 10 cm 1 m 10 m along-track resolution 0 25 50 75 100 range (m) one element full 1.2 m array synthetic aperture 17 mm, flat
Both real apertures degrade in proportion to range; only their starting point differs. The synthetic aperture is flat, because the length it can synthesize grows at exactly the rate that would otherwise blur the image.

The catch is in the word coherent. Summing hundreds of pings only works if the phase of each one is right, and that means knowing where the sonar was at every ping to a small fraction of 5 mm. Underwater there is no GPS (Global Positioning System), and an inertial system drifts well past a millimeter over the tens of seconds an aperture takes to fly. The figure below shows what that costs, on a vertical scale marked in decibels (dB), a ratio measure on which every 20 is a tenfold drop in amplitude.

along-track response of one point target at 50 m 0 dB -10 -20 -30 -40 -200 -100 0 100 200 along-track offset from the target (mm) exact navigation
0.00 mm
nav error, rms
phase error, rms
0.0 dB
peak, vs. perfect
30.0 dB
target above the floor

This is computed in your browser, not drawn: 451 ping positions across the 7.5 m aperture available at 50 m, two-way propagation phase, the element's own beam as an amplitude taper, summed by back-projection the way the beamformer does it. With the navigation exact the measured width is 16.7 mm, which is D/2; the result claimed above is not asserted here, it is measured. The error is modeled as inertial drift, a random walk with the constant offset and the linear ramp removed, because a range bias only shifts the image and a ramp only slides it along-track; what remains is what actually defocuses. Watch what fails first. It is not the width, it is the peak, as energy that belonged in the main lobe leaks out into the floor and buries everything dimmer than the target.

Push the slider to the end, or press discard the phase, and the point collapses to a flat line. That is the whole argument in one picture: with the phase thrown away, 7.5 m of travel and several hundred pings buy exactly nothing, and the target is somewhere in a 7.5 m smear, which is precisely where a single element left it back in step 2. Coherence is not an enhancement layered on top of a synthetic aperture. It is the thing that makes one exist.

So the track has to be recovered from the echoes themselves. Successive pings put different receivers at the same physical point in the water, and those redundant measurements should see an identical patch of seafloor; the delay between them is a direct measurement of how the vehicle actually moved. That is micronavigation. Once the track is known, the summation has to be carried out against true geometry rather than an assumed straight line, which is what time-domain back-projection does.

This is why a synthetic aperture processing chain looks the way it does. The resolution is free in the sense that the physics hands it to you. The price is that you have to measure your own motion to a fraction of a wavelength before you are allowed to collect it.

Same idea, different medium

How SAS differs from SAR.

Every step above is medium-agnostic. A platform moves, it measures a wavefield from a sequence of positions, and the phase relationships between those positions are assembled into an aperture. Radar got there first and synthetic aperture radar is the same construction; the mathematics of image formation is largely shared, and much of the SAS literature is SAR literature with the constants changed.

But the constants differ by five orders of magnitude, and that changes the engineering completely. Sound in water travels 200,000 times slower than light, and almost every difference below follows from that one fact.

Airborne SAR figures are the Sandia MiniSAR / Lynx-class Ka-band reference preset; SAS figures are the 300 kHz array used throughout this page, at 100 m range.
 Airborne SAR, Ka-bandSAS, 300 kHz
Propagation speed300,000 km/s1.5 km/s
Wavelength18 mm5 mm
Fractional bandwidth3.6 %20 %
Working range10 km100 m
Along-track resolution150 mm17 mm
Round trip to max range67 µs133 ms
Platform travel in that time0.4 λ53 λ
Stop-and-hop error, beam edge0.01 λ4 λ
Pulse rate ceiling from range15 kHz7.5 Hz
Position knowledgeGPS and inertial, centimetersno GPS; estimated from the echoes
Propagation speed knowledgeexact for practical purposes1450 to 1550 m/s; must be measured
Polarizationhorizontal and vertical transmit/receive pairs (HH, HV, VH, VV); a primary channelnone; pressure is a scalar field

Sound is slow enough that the platform moves during the pulse

A radar pulse out to 10 km and back takes 67 microseconds, during which a fast aircraft covers under a centimeter. A sonar pulse out to 100 m and back takes 133 milliseconds, during which the vehicle covers 27 cm, more than fifty wavelengths. Transmit and receive genuinely happen from different places.

SAR processing normally assumes they do not, treating the platform as stationary for the duration of each pulse. That approximation, called stop-and-hop, costs a radar about a hundredth of a wavelength of path error at the edge of its beam and costs this sonar about four wavelengths. For SAS it is not an approximation to be tolerated, it is a defect; the geometry has to be modeled with the platform in continuous motion throughout.

Sound is slow enough to cap how fast you can survey

You cannot transmit again until the previous echo is back, so at 100 m range the ping rate cannot exceed 7.5 per second. Step 4 required the synthetic aperture to be sampled every D/2 of travel. Put those together and a single-element sonar could survey at 12 cm/s, which is useless. With 36 receivers laying down 36 samples of the aperture per ping the ceiling becomes 4.5 m/s, which is why the array is there and why survey vehicles run at the speeds they do. A radar's pulse rate ceiling sits in the kilohertz, thousands of times more headroom than it needs.

There is no satellite navigation underwater

A SAR platform knows where it is to a few centimeters, comfortably inside the tolerance step 7 demanded. A submerged vehicle has no such fix, and an inertial system drifts past a millimeter long before an aperture is finished. The track has to be reconstructed from the echoes themselves, which makes micronavigation a load-bearing component of a SAS system rather than a refinement of one. It is the largest single difference in what the two processing chains have to do.

The medium itself is an unknown

Radio waves travel at a speed that can be treated as exact. Sound speed in the sea runs from roughly 1450 to 1550 m/s and varies with temperature, salinity, and depth along the path. Get it wrong by a tenth of one percent and a target at 100 m lands 10 cm away from where it belongs, which is twenty wavelengths of error, so the sound speed profile becomes one more thing to measure or estimate. Two smaller differences compound it: pressure in water is a scalar field, so there is no polarization channel of the kind that carries so much information in radar; and acoustic absorption climbs steeply with frequency, tying resolution and range together far more tightly than in radar. Centimetres are bought by giving up range.

What carries over

More than what does not. Both are coherent, complex-valued imagers that synthesize an aperture out of motion. Both produce single-look complex imagery with fully developed speckle and the same underlying statistics. Both are focused by the same mathematics, whether time-domain back-projection or a wavenumber-domain method. Both show layover, shadow, and range-dependent geometry, and both punish an error in platform position the same way, in phase.

That shared structure is why the SAR literature is the right place to start on a SAS problem, and why a model that has learned to see in one coherent modality is a credible starting point for the other. The differences above are engineering. The physics of coherent imaging is common ground.

ApertureLab simulates this whole chain forward, from a scene file to raw complex time-series, and then focuses it back into an image with the same processing a real system uses. Both of the pages below start from the same physics.