Fix satlas normalization & add MS satlas models (#423)
- fix(models): scale Satlas ResNet native inputs to uint8
Co-Authored-By: Claude Opus 5.5 noreply@anthropic.com
Improve satlas normalization
Add ms satlas
simplify
only reorder for satlas
Add resnet back but within the existing TorchGeoResNetBench class
Fix docs
Co-authored-by: Claude Opus 5.5 noreply@anthropic.com
版权所有:中国计算机学会技术支持:开源发展技术委员会
京ICP备13000930号-9
京公网安备 11010802047560号
torchgeo-bench
A lightweight benchmarking framework for evaluating frozen geospatial foundation models on GeoBench V1/V2 and location encoders on CoordBench. Plug in a backbone or coordinate encoder and run consistent downstream probes through explicit CLI flags and strictly validated Pydantic YAML configuration.
--resumeskips already-computed(dataset, method, model, …)rows. Atomic CSV appends are safe across parallel jobs.contrib_template.py, implement_forward_patch_features, and add a one-file model config. See the Stage 1 guide for a full walkthrough, or the Stage 2 guide to contribute the model back upstream.Installation
For development:
Requires Python 3.12+ and runs on Linux, macOS, and Windows. On Linux x86_64 the install automatically includes GPU-accelerated FAISS KNN (CUDA 12, driver R525+, also works on GPU-less machines); other platforms get CPU FAISS.
Download a dataset
The runner expects datasets under
./data/. To grab GeoBench V1:To download only the datasets needed for a run:
V1 uses the pickle-free JSON-metadata mirror under
data/classification_v1.0_wds/, with a pinned revision and archive SHA-256 verification. Re-download existing pickle-based V1 caches; the readers no longer unpickle metadata.V2 (classification + segmentation) and torchgeo’s EuroSAT, RESISC45, and UC Merced downloaders work the same way (
torchgeo-bench download geobench_v2,torchgeo-bench download eurosat,torchgeo-bench download resisc45,torchgeo-bench download ucmerced). See the documentation for all options.Run a basic experiment
The default device is
cuda:0. On a machine without a working CUDA GPU (or if a GPU run crashes — see troubleshooting), fall back to CPU:Results are appended to
results/models/<model name>.csv, which ship pre-populated with reference results. To start from a clean slate, pass--output results/my_run.csvor setoutput.filein a YAML file passed with--config. Re-run the same command with--resumeto skip completed rows. Each evaluation is saved as soon as it finishes, so a later failure does not discard completed metrics.See
docs/examples/image-run.yamlfor the image configuration fields andrun --config-helpfor the JSON schema. Omitted settings inherit model and dataset defaults; explicit YAML values override them, and supplied flags override YAML. This includes explicitfalse,null, and[].torchgeo-bench,python -m torchgeo_bench, andpython -m torchgeo_bench.cliexpose the same commands. The oldkey=valueand+key=valueoverrides are rejected; migrate scripts to flags or YAML. There is no separate legacy configuration entry point.Measure encoder cost
The standalone
profilecommand measures one fixed real dataset batch and writes JSON to stdout by default. Theflopscommand uses synthetic inputs, does not load dataset samples, and appends compute measurements to a CSV:run,flops,coord, andprofileshare--output-dirfor a common output root and--outputfor an exact file path that takes precedence. Image runs retain theirmodels/,profiles/, andintrinsic_dim/subdirectories. Standalone profiling writesprofile.jsonbeneath an explicit output root.Both accept
--configand--dry-run. Seedocs/examples/profile.yaml,docs/examples/flops.yaml, and the configuration reference. Optionalprofileandintrinsic_dimpasses within an image run retain their separate per-model CSVs unlessoutput.fileexplicitly combines them.Handcrafted classification baseline
The handcrafted model adds deterministic spectral and spatial measurements to the ImageStats baseline. The
handcrafted_level1,handcrafted_level2, andhandcrafted_level3presets select cumulative feature sets;handcrafteddefaults to level 2. Feature width depends on the input bands and available spectral indices.Download the classification datasets and run the level sweep from this checkout:
Future sweeps include ImageStats as level 0 and every classification protocol in the current dataset catalog, including AID, multilabel datasets and EuroSAT’s spatial split. They use all bands with identity input normalization, otherwise retaining the normal 224px resize, KNN-5, validation-selected C, train-plus-validation final refit, and 200 bootstrap draws.
Results go to separate
results/models/handcrafted_level*.csvfiles andimagestats_handcrafted_control.csv. Feature lists and completion status are saved underoutputs/handcrafted/. Resume skips matching completed rows; a missing linear or KNN result is still reported as a failure. Summaries, including--report-only, select only rows matching the current configuration and requested devices, so historical hashes are not mixed into new comparisons.The checked-in CSVs are unchanged historical measurements on 13 protocols from the original handcrafted study, not measurements of the current runner; they contain no AID results. Their recorded numbers and configuration hashes are preserved. Current runs use the current configuration hash and do not treat these historical rows as completed work. Use
--output-dir results/handcrafted-currentto keep a new sweep separate.Use
--levels 1 2,--datasets eurosat resisc45, or--dry-runfor a smaller run. The extractor also works through the normal CLI:All four presets default to all bands and identity normalization. Custom YAML can set
model: {name: handcrafted, kwargs: {level: 3}}; preprocessing belongs underinput, not constructor kwargs.CoordBench — location encoders
torchgeo-bench coordruns the coordinate-only track: point(lon, lat)in, a downstream label out. Benchmarks are streamed directly from the unifiedtaylor-geospatial/coordbenchHuggingFace dataset (PDFM, SatCLIP, SustainBench, CDC PLACES, MOSAIKS/USAVars, the DeepMind/AlphaEarth suite, and more — no local download). A frozen encoder is probed with KNN and a ridge linear head under random or spatial-block cross-validation (regression → R², classification → accuracy).--model mindand--model mind-small(MIND, distilled from AlphaEarth/Climplicit/GeoCLIP/SINR) and--model sincoswork with the base install. The other pretrained encoders (SatCLIP / GeoCLIP / Climplicit / SINR, viarshf) need thecoordbenchextra:pip install "torchgeo-bench[coordbench]", then--model climplicit(etc.). Results land inresults/coordbench_results.csv. Add your own encoder by subclassingLocationEncoder(implement_encode) and settingmodel.targetto its importable Python name, with constructor options undermodel.kwargs. See the CoordBench guide and the runnableFourierLocationEncoderexample.Learn more
Citation
If you use this framework, please cite it (once the
torchgeo-benchpaper is available):License
MIT.