项目文件夹

文件
wehub-resource-sync 593b94c120
pytest / Unit Tests (push) Has been cancelled
pytest / Integration (integration_tests_a) (push) Has been cancelled
pytest / Integration (integration_tests_b) (push) Has been cancelled
pytest / Integration (integration_tests_c) (push) Has been cancelled
pytest / Integration (integration_tests_d) (push) Has been cancelled
pytest / Integration (integration_tests_e) (push) Has been cancelled
pytest / Integration (integration_tests_f) (push) Has been cancelled
pytest / Integration (integration_tests_g) (push) Has been cancelled
pytest / Integration (integration_tests_h) (push) Has been cancelled
pytest / Integration (integration_tests_i) (push) Has been cancelled
pytest / Integration (integration_tests_j) (push) Has been cancelled
pytest / Distributed (distributed_a) (push) Has been cancelled
pytest / Distributed (distributed_b) (push) Has been cancelled
pytest / Distributed (distributed_c) (push) Has been cancelled
pytest / Distributed (distributed_d) (push) Has been cancelled
pytest / Distributed (distributed_e) (push) Has been cancelled
pytest / Distributed (distributed_f) (push) Has been cancelled
pytest / Minimal Install (push) Has been cancelled
pytest / Event File (push) Has been cancelled
pytest (slow) / py-slow (push) Has been cancelled
Publish JSON Schema / publish-schema (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:49:20 +08:00

113 行
4.0 KiB
Markdown

此文件含有模棱两可的 Unicode 字符
此文件含有可能会与其他字符混淆的 Unicode 字符。 如果您是想特意这样的,可以安全地忽略该警告。 使用 Escape 按钮显示他们。
# Semantic Segmentation: UNet, SegFormer, and FPN
[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/ludwig-ai/ludwig/blob/main/examples/semantic_segmentation/semantic_segmentation.ipynb)
Semantic segmentation assigns a class label to every pixel in an image.
This example trains three different decoder architectures on the **CamSeq01**
urban driving dataset (101 images, 32 semantic classes) and compares their
accuracy/speed trade-offs.
## Decoder comparison
| Decoder | Architecture | Recommended encoder | Approx. extra params | Best for |
| ----------- | -------------------------------------------------------------------------- | ---------------------------------- | ------------------------ | ------------------------------------------------------------------------ |
| `unet` | Symmetric encoder-decoder with skip connections; configurable `num_stages` | Built-in `unet` encoder | ~31M (depth 4) | General purpose baseline, no pretrained backbone needed |
| `segformer` | Lightweight all-MLP head fusing multi-scale ViT features | `dinov2` (DINOv2-base, pretrained) | ~2M head + ~86M backbone | Highest accuracy; transformer features transfer well to dense prediction |
| `fpn` | Feature Pyramid Network top-down pathway with lateral connections | `efficientnet` (pretrained) | ~2M head + ~5M backbone | Fast inference; handles objects at multiple scales efficiently |
## Dataset
[CamSeq01](https://mi.eng.cam.ac.uk/research/projects/VideoRec/CamSeq01/) is a
set of 101 road-scene images captured in Cambridge, UK at 960×720 resolution
with 32 semantic class annotations.
Ludwig ships a built-in downloader — see [`camseq.py`](camseq.py) for the
standalone script or use `from ludwig.datasets import camseq` in Python.
## Config files
| File | Decoder | Notes |
| ------------------------ | ----------- | ---------------------------------------- |
| `config_camseq.yaml` | `unet` | Original baseline config |
| `config_unet_depth.yaml` | `unet` | Shows the `num_stages` parameter |
| `config_segformer.yaml` | `segformer` | DINOv2 backbone, fine-tuned end-to-end |
| `config_fpn.yaml` | `fpn` | EfficientNet backbone, larger batch size |
## Running the examples
**Prerequisites**: a CUDA-capable GPU. An A100 or equivalent is recommended
for the SegFormer run; the UNet and FPN configs run well on a single V100/3090.
```bash
pip install 'ludwig[vision]'
```
### UNet (configurable depth)
```bash
python camseq.py # uses config_camseq.yaml (depth 4 by default)
```
Or with the explicit depth config:
```bash
ludwig train --config config_unet_depth.yaml
```
### SegFormer
```bash
ludwig train --config config_segformer.yaml
```
### FPN
```bash
ludwig train --config config_fpn.yaml
```
### UNet depth ablation
```bash
python unet_depth_sweep.py
```
This script trains models with `num_stages` ∈ {2, 3, 4, 5} and prints a
summary table of parameter count vs. best validation loss vs. training time.
### Interactive notebook
Open `semantic_segmentation.ipynb` locally or click the Colab badge above.
The notebook walks through all three decoders and produces side-by-side
visualisations of their predictions.
## Key config parameters
### UNet decoder
```yaml
decoder:
type: unet
num_stages: 4 # 2–5; input size must be divisible by 2^num_stages
num_fc_layers: 0
conv_norm: batch
```
### SegFormer decoder
```yaml
decoder:
type: segformer
hidden_size: 256 # MLP projection width
dropout: 0.1
```
### FPN decoder
```yaml
decoder:
type: fpn
num_channels: 256 # lateral projection width at each pyramid level
num_levels: 4 # number of pyramid levels (typical range 2–5)
```