|
autonomy_benchmarks v1.0.0
|
This repository supports the following benchmarks for object detection in automated driving systems:
The nuScenes dataset is not redistributed here. You must download it under the nuScenes Terms of Use and make it available via the companion autonomy_datasets package.
Supported Datasets:
This benchmark uses nuScenes Dataset via the autonomy_datasets ROS package.
Metrics are computed based on the following assumptions:
Ground-truth labels carry the full nuScenes annotation category (e.g. vehicle.car); predictions already carry the detection class name. Categories are grouped into the 10 evaluated detection classes following the official general_to_detection mapping, categories outside those classes are dropped and not evaluated, and static_object.bicycle_rack is retained (as bike_rack) solely for the bike-rack filter below and is never scored.
vehicle.car: carvehicle.truck: truckvehicle.bus.bendy: busvehicle.bus.rigid: busvehicle.trailer: trailervehicle.construction: construction_vehiclehuman.pedestrian.adult: pedestrianhuman.pedestrian.child: pedestrianhuman.pedestrian.construction_worker: pedestrianhuman.pedestrian.police_officer: pedestrianvehicle.motorcycle: motorcyclevehicle.bicycle: bicyclemovable_object.trafficcone: traffic_conemovable_object.barrier: barrierstatic_object.bicycle_rack: bike_rack (filter only, not scored)animal: droppedhuman.pedestrian.personal_mobility: droppedhuman.pedestrian.stroller: droppedhuman.pedestrian.wheelchair: droppedmovable_object.debris: droppedmovable_object.pushable_pullable: droppedvehicle.emergency.ambulance: droppedvehicle.emergency.police: droppedLabels and predictions are only considered if they fall into a class-specific detection range.
<= 30 m<= 30 m<= 40 m<= 40 m<= 40 m<= 50 m<= 50 m<= 50 m<= 50 m<= 50 m{0.5, 1.0, 2.0, 4.0} meters. For each match threshold, average precision (ap) is calculated by integrating the recall-precision curve for recalls and precisions > 0.1 (points at or below either threshold are excluded). The mean average precision (map) is the average over match thresholds and classes.ate, ...) are calculated using a match threshold of 2.0 m.| Metric | Description |
|---|---|
ap_0.5_{barrier,traffic_cone,...} | Average precision for class with maximum match distance of 0.5 meters. |
ap_1.0_{barrier,traffic_cone,...} | Average precision for class with maximum match distance of 1 meter. |
ap_2.0_{barrier,traffic_cone,...} | Average precision for class with maximum match distance of 2 meters. |
ap_4.0_{barrier,traffic_cone,...} | Average precision for class with maximum match distance of 4 meters. |
map_{barrier,traffic_cone,...} | Mean average precision for class over all match distance thresholds. |
map | Mean average precision all match distance thresholds and all classes. |
ate_2.0_{barrier,traffic_cone,...} | Average translation error for class as Euclidean distance in meters. |
mate_2.0 | Mean average translation error over all classes. |
ase_2.0_{barrier,traffic_cone,...} | Average scale error for class as 1 - IoU after aligning centers and orientation. |
mase_2.0 | Mean average scale error over all classes. |
aoe_2.0_{barrier,traffic_cone,...} | Average orientation error for class as smallest yaw angle difference between prediction and ground truth in radians. Orientation errors for traffic_cones are ignored and barriers are only evaluated up to 180 degrees. |
maoe_2.0 | Mean average orientation error over all classes. |
ave_2.0_{barrier,traffic_cone,...} | Average velocity error for class as absolute velocity error in m/s. Velocities for barriers and traffic_cones are ignored. |
mave_2.0 | Mean average velocity error over all classes. |
aae_2.0_{barrier,traffic_cone,...} | Average attribute error for class as 1 - attribute_accuracy. Attribute errors for barriers and traffic_cones are ignored. Returns null when the dataset loader does not provide attribute annotations. |
maae_2.0 | Mean average attribute error over all classes. In the current implementation this is always a numeric value (defaults to 1.0 if no valid attribute-error values are available). |
nds | nuScenes Detection Score (NDS) combining map and five TP error scores with weights 5-1-1-1-1-1, normalised by 10: NDS = (5·mAP + max(1−mATE,0) + max(1−mASE,0) + max(1−mAOE,0) + max(1−mAVE,0) + max(1−mAAE,0)) / 10. In the current implementation the denominator is always 10. |
To contribute a new benchmark for a dataset or evaluation protocol:
AutonomyBenchmark.required_inputs(), compute_sample_metrics(), and compute_aggregated_metrics().__init__ (thresholds, per-class ranges, and metric rules as instance attributes); keep static lookup tables (e.g. category-to-class mappings) as module-level _CONSTANT_NAME constants.benchmark:=<name> launch argument.