Engineering Background Removal#
How the withoutBG models are trained and measured. A model change ships only when it scores better on held-out evaluation data. Results are published, including the comparisons withoutBG loses.
The problem#
- Hair, fur, semi-transparent objects (glass, veils, smoke) and motion blur need a soft alpha matte, not a hard mask.
- Low contrast between subject and background makes edges ambiguous.
- Models that score well on academic test sets often fail on real photos: product shots, phone pictures, mixed lighting.
Two models#
| Open weights | API | |
|---|---|---|
| Runs on | Your CPU or GPU (ONNX), or Core ML in the Mac app | AWS, Frankfurt |
| Distributed as | Python package, Docker image, CLI, ONNX file | REST API, web app, plugins |
| Images leave your machine | No | Yes, processed in memory and not stored |
| License | Apache 2.0 + Meta DINOv3 License | Closed, paid per image |
The open-weights model is a single ONNX graph (~495 MB) that runs depth estimation, segmentation, matting and refinement in one pass, at up to 768 px output. It is built on a DINOv3 backbone. The API model is larger and trained on the same data pipeline. See the open weights vs API comparison.
Training data#
The dataset is the hardest and most expensive part of the work.
- Studio captures. Real photos shot against known backgrounds, so the true matte can be recovered (how the dataset is made).
- Harmonized composites. A subject pasted onto a new background has mismatched light and color, and a model learns to spot those edges. A GAN adjusts lighting and color so composites look like real photos.
- 3D renders. Blender scenes with randomized lights, materials and cameras add variety that flips and crops can't.
- Realistic compositing. Backgrounds are varied and objects get contact shadows.
- Human annotation and review. Paid annotators label hard cases. Generated samples are reviewed, and bad edges or wrong labels are removed before training.
Training#
- Experiments start small on local hardware and scale up on cloud GPUs.
- Curriculum: small, easier images first, then larger and harder ones.
- Large images are cropped, with crops sampled around edges and fine detail so training memory goes where the errors are.
- 300+ experiments so far, each one logged with its data, settings and metrics in Weights & Biases.
Evaluation#
Predicted mattes are compared to ground-truth mattes with the standard alpha matting metrics. Lower is better for all of them.
- SAD (sum of absolute differences): total error over the image.
- MSE (mean squared error): punishes large errors more than small ones.
- Gradient error: blurry or jagged edges, lost hair detail.
- Connectivity error: holes and floating fragments in the matte.
Definitions and reference scores: alpha matting evaluation benchmark. Evaluation sets: the public withoutBG100 and the benchmark set used on the compare pages. Results are read per category as well as overall, because a small average gain can hide a loss on hair or transparent objects.
Published results#
| Open weights vs | MAE | Gradient error | Connectivity |
|---|---|---|---|
| BRIA RMBG | 0.033 vs 0.023 worse | 0.039 vs 0.052 better | 0.032 vs 0.022 worse |
| BiRefNet General | 0.033 vs 0.033 better | 0.039 vs 0.054 better | 0.032 vs 0.033 better |
| BiRefNet Lite | 0.033 vs 0.047 better | 0.039 vs 0.057 better | 0.032 vs 0.045 better |
| IS-Net | 0.033 vs 0.064 better | 0.039 vs 0.070 better | 0.032 vs 0.063 better |
| u2net | 0.033 vs 0.084 better | 0.039 vs 0.087 better | 0.032 vs 0.084 better |
| silueta | 0.033 vs 0.089 better | 0.039 vs 0.094 better | 0.032 vs 0.089 better |
| u2netp | 0.033 vs 0.092 better | 0.039 vs 0.088 better | 0.032 vs 0.091 better |
Some rows say “worse”. They are published anyway. Every comparison, including the API model against remove.bg, is in Compare. Result galleries show hard cases and failures: open weights, API.
Inference#
- API: runs on AWS in Frankfurt. An EU region keeps processing under GDPR, and most customers are in Europe. Images are processed in memory and not stored.
- Web app: the browser sends a copy of at most 1024 px, gets the alpha matte back and applies it to the full-resolution original locally. The original never leaves the device.
- Open weights: ONNX on CPU or GPU through the Python package, Docker or CLI; Core ML in the Mac app.
Licensing#
Open weights: Apache 2.0 for the withoutBG parts, Meta DINOv3 License for the DINOv3 backbone weights. Details: license.
Contribute#
Failure cases make the next model better. Issues, pull requests and hard images are welcome on GitHub. Models: Hugging Face. Images: Docker Hub.