Teach a model the seafloor before it ever goes to sea.

ApertureLab is a synthetic aperture sonar simulation and beamforming workbench; built to accelerate the AI systems that will map and understand the ocean floor at scale.

Run the beamformer in your browser

A real back-projection beamformer, 564 pings over a 9.4 m track, running in the page. Move the aperture slider and watch the scene come into focus.

↓

What ApertureLab unlocks.

SAR foundation models exist because satellite imagery was cheap, abundant, and labeled. Sonar data requires ships, dive operations, and restricted access. ApertureLab generates physics-accurate labeled sonar data at satellite-imagery scale; the missing precondition that makes the following possible for the first time.

Synthetic Datasets

Labeled imagery, by construction

Every simulated image carries ground-truth labels derived automatically from the scene file: object class, position, orientation, burial depth, seafloor type, shadow mask. Six annotation categories, written out with the image in COCO format, with the pixel size and scene extent attached so a georeference can follow.

  • Oriented object footprints from real geometry and heading
  • Buffered polylines for cables and pipeline runs
  • Traced outlines for every seafloor type in the scene
  • Full domain variation: depth, range, bottom type, aspect, burial

A scene is a file, so the dataset is a parameter sweep. The cost of the next thousand images is the cost of the first one, times a thousand, with no ship.

Foundation Models

No general-purpose foundation model has been trained on SAS imagery. Yet.

Not because the architecture does not exist (ViT, MAE, and contrastive pretraining are all mature), but because training data at the necessary scale never existed. ApertureLab closes that gap with parametric scenes spanning the full space of acoustic environments.

  • Sim-to-real transfer for autonomous underwater vehicle perception
  • Environment-agnostic feature representations across seabed types
  • Edge-case and rare-target coverage unavailable in field collections
  • The sonar analogue to ImageNet-scale pretraining

This is the next rung. The ladder already exists.

Vision-Language Models

VLMs already respond to sonar. Fine-tuning is the obvious next step.

The 2026 IGARSS work shows VLMs classify SAS targets at 0.946 AUC using only a text prompt describing highlight-shadow geometry, with zero domain-specific training. Scene captions generated at render time give the fine-tuning step what it needs: captioned imagery whose every claim about the scene is true by construction.

  • Auto-generated natural-language captions paired with every image
  • Query sonar archives by English phrase
  • Caption imagery for non-expert operators and analysts
  • A path to zero-shot object recognition across unseen target types

The zero-shot baseline is 0.946 AUC. Domain fine-tuning is the straightforward next step.

Evidence

From one English sentence to a labeled dataset.

The entire specification, written through Claude Code “create a 60 by 100 m scene with a grid of different seafloor backgrounds and a target in the middle of each one, so I can crop 256 by 256 chips of the same object on different bottoms for ML training.”

This is what came back, rendered overnight on one workstation, physics through the entire chain: sixty seafloor targets on ten varied bottom classes (clean sand, mud, rock, gravel, coarse gravel, silt, rippled sand, two rocky-sediment variants, and shelly sand hash), with roughness and reflectivity jittered per cell so no two cells match. Hover over the image to reveal the labels; every green box, name, and seafloor zone outline is generated automatically from the scene file at simulation time, not hand-annotated.

The same physics that renders a labeled scene forward is what trains the model that reads one backward.

Evidence

Every image arrives already labeled.

The same scene, twice. On the left, the sonar image the physics produces. On the right, the object footprints and seafloor-type boundaries that come with it. Nothing here was annotated by hand; the labels are emitted from the scene file at simulation time, so they are exact by construction rather than accurate to within a human's patience.

A simulated synthetic aperture sonar scene beside the same scene overlaid with object bounding boxes and traced seafloor-zone outlines
Object class, position, orientation, burial depth, and the traced outline of each seafloor type, written out with the image in COCO format.

The whole thing

An empty window to a labeled sonar image.

The whole build, sped up: an empty window to a focused SAS image over 100 m of range by 60 m of along-track, with a pixel-accurate label on every object in it. The pages below take apart what you just watched.

Explore

Where to go deeper.

Each page below walks through one capability and ends in the imagery it produces.

Isaac D. Gerg, Ph.D.

AI Scientist, ClimateAI · gergltd.com

Most underwater AI projects fail at the same seam: the ML team does not understand the acoustics, and the sonar team does not understand deep learning. I work at both ends. Twenty years of research spanning the complete signal chain, from acoustic wave scattering and IQ time-series recording, through beamforming and image formation, to deep embeddings for detection, segmentation, and compression.

ApertureLab is what that understanding looks like as a piece of software. If you are building AI systems that need to understand the physical world, I would like to talk.