Engineering Background Removal#

How the withoutBG models are trained and measured. A model change ships only when it scores better on held-out evaluation data. Results are published, including the comparisons withoutBG loses.

The problem#

  • Hair, fur, semi-transparent objects (glass, veils, smoke) and motion blur need a soft alpha matte, not a hard mask.
  • Low contrast between subject and background makes edges ambiguous.
  • Models that score well on academic test sets often fail on real photos: product shots, phone pictures, mixed lighting.

Two models#

Open weightsAPI
Runs onYour CPU or GPU (ONNX), or Core ML in the Mac appAWS, Frankfurt
Distributed asPython package, Docker image, CLI, ONNX fileREST API, web app, plugins
Images leave your machineNoYes, processed in memory and not stored
LicenseApache 2.0 + Meta DINOv3 LicenseClosed, paid per image

The open-weights model is a single ONNX graph (~495 MB) that runs depth estimation, segmentation, matting and refinement in one pass, at up to 768 px output. It is built on a DINOv3 backbone. The API model is larger and trained on the same data pipeline. See the open weights vs API comparison.

Training data#

The dataset is the hardest and most expensive part of the work.

  • Studio captures. Real photos shot against known backgrounds, so the true matte can be recovered (how the dataset is made).
  • Harmonized composites. A subject pasted onto a new background has mismatched light and color, and a model learns to spot those edges. A GAN adjusts lighting and color so composites look like real photos.
  • 3D renders. Blender scenes with randomized lights, materials and cameras add variety that flips and crops can't.
  • Realistic compositing. Backgrounds are varied and objects get contact shadows.
  • Human annotation and review. Paid annotators label hard cases. Generated samples are reviewed, and bad edges or wrong labels are removed before training.

Training#

  • Experiments start small on local hardware and scale up on cloud GPUs.
  • Curriculum: small, easier images first, then larger and harder ones.
  • Large images are cropped, with crops sampled around edges and fine detail so training memory goes where the errors are.
  • 300+ experiments so far, each one logged with its data, settings and metrics in Weights & Biases.

Evaluation#

Predicted mattes are compared to ground-truth mattes with the standard alpha matting metrics. Lower is better for all of them.

  • SAD (sum of absolute differences): total error over the image.
  • MSE (mean squared error): punishes large errors more than small ones.
  • Gradient error: blurry or jagged edges, lost hair detail.
  • Connectivity error: holes and floating fragments in the matte.

Definitions and reference scores: alpha matting evaluation benchmark. Evaluation sets: the public withoutBG100 and the benchmark set used on the compare pages. Results are read per category as well as overall, because a small average gain can hide a loss on hair or transparent objects.

Published results#

Open weights vsMAEGradient errorConnectivity
BRIA RMBG0.033 vs 0.023 worse0.039 vs 0.052 better0.032 vs 0.022 worse
BiRefNet General0.033 vs 0.033 better0.039 vs 0.054 better0.032 vs 0.033 better
BiRefNet Lite0.033 vs 0.047 better0.039 vs 0.057 better0.032 vs 0.045 better
IS-Net0.033 vs 0.064 better0.039 vs 0.070 better0.032 vs 0.063 better
u2net0.033 vs 0.084 better0.039 vs 0.087 better0.032 vs 0.084 better
silueta0.033 vs 0.089 better0.039 vs 0.094 better0.032 vs 0.089 better
u2netp0.033 vs 0.092 better0.039 vs 0.088 better0.032 vs 0.091 better
Mean error against ground-truth alpha mattes on Bench v0 (36 images, same inputs for every model). Lower is better. Each row links to the side-by-side images.

Some rows say “worse”. They are published anyway. Every comparison, including the API model against remove.bg, is in Compare. Result galleries show hard cases and failures: open weights, API.

Inference#

  • API: runs on AWS in Frankfurt. An EU region keeps processing under GDPR, and most customers are in Europe. Images are processed in memory and not stored.
  • Web app: the browser sends a copy of at most 1024 px, gets the alpha matte back and applies it to the full-resolution original locally. The original never leaves the device.
  • Open weights: ONNX on CPU or GPU through the Python package, Docker or CLI; Core ML in the Mac app.

Licensing#

Open weights: Apache 2.0 for the withoutBG parts, Meta DINOv3 License for the DINOv3 backbone weights. Details: license.

Contribute#

Failure cases make the next model better. Issues, pull requests and hard images are welcome on GitHub. Models: Hugging Face. Images: Docker Hub.