autonomy_datasets v1.6.0
Loading...
Searching...
No Matches
Implementation Details

Supported Datasets

This repository supports various automated driving datasets including:

  • **Waymo Open Dataset**
  • **nuScenes**
  • **MAN TruckScenes**
  • **NVIDIA PhysicalAI AV Dataset**
  • **DrivIng**
  • **TUM Traffic**
  • **Zenseact Open Dataset**
  • **FZI-AURA**
  • **Thinking Cars Datasets** available on request for commercial use and custom data
  • **Contributions** adding more open datasets are welcome

Waymo Open Dataset

Waymo Waymo Open Dataset

Rviz Screenshot Waymo Open Dataset

Split Samples
all 198.068
training 158.081
validation 39.987
Source Topic Type Description
Sensor: Top Lidar /lidar_01/point_cloud sensor_msgs/msg/PointCloud2 Raw sensor data from top lidar as point cloud with fields (x, y, z, intensity, elongation) in sensor frame.
Sensor: Front Camera /camera_01/image_raw/camera_01/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=1280px, width=1920px) from front camera.
Sensor: Front-Left Camera /camera_02/image_raw/camera_02/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=1280px, width=1920px) from front-left camera.
Sensor: Front-Right Camera /camera_03/image_raw/camera_03/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=1280px, width=1920px) from front-right camera.
Sensor: Side-Left Camera /camera_04/image_raw/camera_04/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=886px, width=1920px) from side-left camera.
Sensor: Side-Right Camera /camera_05/image_raw/camera_05/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=886px, width=1920px) from side-right camera.
EgoData /ego_data perception_msgs/msg/EgoData Ego-vehicle's dimensions and dynamics state in map frame.
Annotation: 3D Lidar Objects /object_list/lidar_01 perception_msgs/msg/ObjectList Annotated 3D objects (HEXAMOTION model) in vehicle frame. Default: Only objects with min. 1 point in top lidar point cloud.
Annotation: 2D Camera Objects /object_list/cameras perception_msgs/msg/ObjectList Annotated 2D objects (CAMERA2D model) in camera frame. Note: Currently no visualization is shown for this data type in RViz.
Meta Information: Object Annotations /object_list/lidar_01/meta_info/object_list/camera_01/meta_info/object_list/camera_all/meta_info autonomy_datasets_msgs/msg/ObjectListMetaInfo Annotations without a representation in perception_msgs/msg/Object: original_class, num_lidar_pts and difficulty_level. Associated with the object list via the header stamp and the object id.
Transformations /tf, /tf_static tf2_msgs/msg/TFMessage Static transformations to all sensor frames and dynamic transformation from map to vehicle frame.

Usage

Download the dataset and ensure the following folder structure is correct:

$DATASET_DIR/
waymo_open_dataset/
training/
camera_box/
*.parquet
...
...
validation/
camera_box/
*.parquet
...
...

Run the ROS node to convert and store the data to rosbags while visualizing it in Rviz.

ros2 launch autonomy_datasets autonomy_datasets.launch.py dataset:=waymo_open_dataset

nuScenes Dataset

nuScenes nuScenes

Rviz Screenshot nuScenes Dataset

Split Samples
training 28.130
validation 6.019
training_mini 25
validation_mini 17
Source Topic Type Description
Sensor: Top Lidar (Velodyne HDL-32E) /lidar_01/point_cloud sensor_msgs/msg/PointCloud2 Raw sensor data from top lidar as point cloud with fields (x, y, z, intensity, timestamp).
Sensor: Front Camera (Basler acA1600-60gc) /camera_01/image_raw/camera_01/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=900px, width=1600px) from front camera.
Sensor: Front-Right Camera (Basler acA1600-60gc) /camera_02/image_raw/camera_02/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=900px, width=1600px) from front-right camera.
Sensor: Back-Right Camera (Basler acA1600-60gc) /camera_03/image_raw/camera_03/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=900px, width=1600px) from back-right camera.
Sensor: Back Camera (Basler acA1600-60gc) /camera_04/image_raw/camera_04/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=900px, width=1600px) from back camera.
Sensor: Back-Left Camera (Basler acA1600-60gc) /camera_05/image_raw/camera_05/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=900px, width=1600px) from back-left camera.
Sensor: Front-Left Camera (Basler acA1600-60gc) /camera_06/image_raw/camera_06/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=900px, width=1600px) from front-left camera.
Sensor: Front Radar (Continental ARS 408-21) /radar_01/point_cloud sensor_msgs/msg/PointCloud2 Radar detections from front radar as point cloud with fields (x, y, z, radial_velocity, rcs).
Sensor: Front-Right Radar (Continental ARS 408-21) /radar_02/point_cloud sensor_msgs/msg/PointCloud2 Radar detections from front-right radar as point cloud with fields (x, y, z, radial_velocity, rcs).
Sensor: Back-Right Radar (Continental ARS 408-21) /radar_03/point_cloud sensor_msgs/msg/PointCloud2 Radar detections from back-right radar as point cloud with fields (x, y, z, radial_velocity, rcs).
Sensor: Back-Left Radar (Continental ARS 408-21) /radar_04/point_cloud sensor_msgs/msg/PointCloud2 Radar detections from back-left radar as point cloud with fields (x, y, z, radial_velocity, rcs).
Sensor: Front-Left Radar (Continental ARS 408-21) /radar_05/point_cloud sensor_msgs/msg/PointCloud2 Radar detections from front-left radar as point cloud with fields (x, y, z, radial_velocity, rcs).
EgoData /ego_data perception_msgs/msg/EgoData Ego-vehicle's dimensions and dynamics state (EGO model) in map frame. Pose from the dataset's ego poses, velocity, acceleration, yaw rate, steering angle, standstill flag, turn indicator and brake light from the CAN bus expansion.
Annotation: 3D Lidar Objects /object_list/lidar_01 perception_msgs/msg/ObjectList Annotated 3D objects (HEXAMOTION model) visible in lidar scan. As nuScenes stores no object dynamics, the absolute velocity, acceleration and yaw rate are differentiated from the annotation positions and yaw angles of the neighboring keyframes.
Annotation: 3D Front Camera Objects /object_list/camera_01 perception_msgs/msg/ObjectList Annotated 3D objects (HEXAMOTION model) visible in front camera image. As nuScenes stores no object dynamics, the absolute velocity, acceleration and yaw rate are differentiated from the annotation positions and yaw angles of the neighboring keyframes.
Transformations /tf, /tf_static tf2_msgs/msg/TFMessage Static transformations to all sensor frames and dynamic transformation from map to vehicle frame.
Detection: Detected 3D Objects /object_list/detected perception_msgs/msg/ObjectList Detected 3D objects (HEXAMOTION model) from Megvii baseline.
Meta Information: Object Annotations /object_list/lidar_01/meta_info/object_list/camera_01/meta_info/object_list/detected/meta_info autonomy_datasets_msgs/msg/ObjectListMetaInfo Annotations without a representation in perception_msgs/msg/Object: original_class, num_lidar_pts, num_radar_pts, num_points, attribute and detection_score. Associated with the object list via the header stamp and the object id.
Map map_contents (parameter) string Lanelet2 map (OSM XML) of the current scene's location, converted from the nuScenes map expansion. Updated on every scene change, analogous to lanelet2_map_server.

The Lanelet2 conversion is controlled via the publish_lanelet2_map (enable/disable) and nuscenes_lanelet2_lane_width (assumed lane width in meters) parameters. Lanes and lane connectors are converted to road lanelets (boundaries synthesized by offsetting the centerline by half the lane width), and pedestrian crossings to crosswalk lanelets. Conversion requires the nuScenes map-expansion data under maps/expansion/.

The converted map is stored next to the rosbag data of the scene it belongs to: the map as map.osm and its origin as map.yaml inside the scene's rosbag directory. Replaying a rosbag restores the map from there instead of converting it again, unless publish_lanelet2_map is disabled.

Usage

Download the dataset (including CAN Bus and Map Expansion) and ensure the following folder structure is correct:

$DATASET_DIR/
nuscenes/
can_bus/
detection-megvii/ # optional
megvii_*.json
maps/
basemap/
*.png
expansion/
*.json
prediciton/
prediction_scenes.json
*.png
...
samples/
CAM_BACK/
*.jpg
...
sweeps/
CAM_BACK/
*.jpg
...
v1.0-mini/
*.json
v1.0-test/
*.json
v1.0-trainval/
*.json

Run the ROS node to convert and store the data to rosbags while visualizing it in Rviz.

ros2 launch autonomy_datasets autonomy_datasets.launch.py dataset:=nuscenes

MAN TruckScenes Dataset

CC BY-NC-SA AWS Open Data

Rviz Screenshot MAN TruckScenes Dataset

MAN TruckScenes is a public dataset recorded from a heavy truck. It comprises 747 scenes of 20 seconds each, annotated at 2 Hz, recorded with 6 lidars, 6 radars, 4 cameras and a high-precision GNSS. The dataset reuses the nuScenes database schema and is licensed under CC BY-NC-SA 4.0.

Split Scenes Samples
train 523 approx. 20.900
val 75 approx. 3.000
test 149 approx. 5.900
mini_train 8 approx. 320
mini_val 2 approx. 80

‍The test split is released without object annotations, so the object list topics are published empty for it.

Source Topic Type Description
Sensor: 6 Lidars /lidar_01/point_cloud ... /lidar_06/point_cloud sensor_msgs/msg/PointCloud2 Point clouds in the respective sensor frame with float32 fields (x, y, z, intensity) and a float64 absolute-seconds timestamp, preserving native per-point timing. Topic order is left, right, top-front, top-left, top-right, and rear.
Sensor: 6 Radars /radar_01/point_cloud ... /radar_06/point_cloud sensor_msgs/msg/PointCloud2 Radar detections in the respective sensor frame with float32 fields (x, y, z, vrel_x, vrel_y, vrel_z, rcs). Topic order is left-front, right-front, right-side, right-back, left-back, and left-side.
Sensor: 4 Cameras /camera_01/image_raw ... /camera_04/image_raw/camera_01/camera_info ... /camera_04/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Undistorted and rectified RGB images with native calibration; topic order is left-front, right-front, right-back, and left-back.
EgoData /ego_data perception_msgs/msg/EgoData Ego-vehicle's dimensions and dynamics state in the UTM-WGS84 (zone U32) map frame, with velocities, accelerations, and yaw rate taken from the native ego_motion_chassis table.
Annotation: 3D Lidar Objects /object_list/lidar_01 perception_msgs/msg/ObjectList Annotated 3D objects (HEXAMOTION model) in the left lidar frame. Default: Only objects with min. 1 point in the lidar point cloud.
Annotation: 3D Camera Objects /object_list/camera_01 perception_msgs/msg/ObjectList Annotated 3D objects (HEXAMOTION model) visible in the left front camera image.
Meta Information: Object Annotations /object_list/lidar_01/meta_info/object_list/camera_01/meta_info autonomy_datasets_msgs/msg/ObjectListMetaInfo Annotations without a representation in perception_msgs/msg/Object: original_class, num_lidar_pts, num_radar_pts, num_points and attribute. Associated with the object list via the header stamp and the object id.
Transformations /tf, /tf_static tf2_msgs/msg/TFMessage Static transformations to all sensor frames and dynamic transformation from map to vehicle frame.

Usage

Select the split using dataset_split in params_truckscenes.yml. Missing data is downloaded automatically from the AWS Open Data registry without requiring credentials; alternatively download the archives manually and unpack them into the following folder structure:

$DATASET_DIR/
truckscenes/
samples/
CAMERA_LEFT_FRONT/
*.jpg
LIDAR_LEFT/
*.pcd
RADAR_LEFT_FRONT/
*.pcd
...
v1.2-mini/
*.json
v1.2-test/
*.json
v1.2-trainval/
*.json

‍Only keyframe sensor data (samples/) is downloaded, because the adapter publishes annotated keyframes only. The unannotated sweeps/ are skipped to keep the required disk space low.

Run the ROS node to download, convert, and store the data to rosbags while visualizing it in Rviz.

ros2 launch autonomy_datasets autonomy_datasets.launch.py dataset:=truckscenes

NVIDIA PhysicalAI AV Dataset

NVIDIA License Hugging Face

Rviz Screenshot NVIDIA PhysicalAI AV Dataset

The number of samples depends on the configurable selected sensor modalities:

| Sensor Modalities | Sensor Setup | Samples | | ---— | ---— | -— | | Camera | 7 cameras at 30 Hz | 306.152 (20 seconds each) | 183.691.200 | | Camera + Lidar | 7 cameras + 360 deg lidar at 10 Hz | 298.326 (20 seconds each) | 59.665.200 | | Camera + Radar | 7 camera + up to 10 radars at 10 Hz | 160.761 (20 seconds each) | 32.152.200 | | Camera + Lidar + Radar | 7 camera + 360 deg lidar at 10 Hz + up to 10 radars at 10 Hz | TODO (20 seconds each) | TODO |

The provided default splits contain only samples including all sensor modalities (Camera + Lidar + Radar).

Split Country Scenes Samples
all All 85.082 approx. 17.016.400
all Germany 7.247 approx. 1.449.400
train Germany 3.694 approx. 738.800
valid Germany 2.044 approx. 408.800
test Germany 1.509 approx. 301.800
Source Topic Type Description
Sensor: Top Lidar /lidar_01/point_cloud sensor_msgs/msg/PointCloud2 Raw sensor data from top lidar as point cloud with fields (x, y, z, intensity) in sensor frame.
Sensor: Front Tele Camera (30° FOV) /camera_01/image_raw/camera_01/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=1080px, width=1920px) from front tele camera.
Sensor: Front Wide Camera (120° FOV) /camera_02/image_raw/camera_02/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=1080px, width=1920px) from front wide camera.
Sensor: Left Cross Camera (120° FOV) /camera_03/image_raw/camera_03/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=1080px, width=1920px) from left cross camera.
Sensor: Right Cross Camera (120° FOV) /camera_04/image_raw/camera_04/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=1080px, width=1920px) from right cross camera.
Sensor: Rear-Left Camera (70° FOV) /camera_05/image_raw/camera_05/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=1080px, width=1920px) from rear-left camera.
Sensor: Rear-Right Camera (70° FOV) /camera_06/image_raw/camera_06/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=1080px, width=1920px) from rear-right camera.
Sensor: Rear Tele Camera (30° FOV) /camera_07/image_raw/camera_07/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Raw RGB images (height=1080px, width=1920px) from rear tele camera.
EgoData /ego_data perception_msgs/msg/EgoData Ego-vehicle's dimensions and dynamics state in map frame.
Annotation: 3D Lidar Objects /object_list/lidar_01 perception_msgs/msg/ObjectList Annotated 3D objects (HEXAMOTION model) in vehicle frame.
Meta Information: Object Annotations /object_list/lidar_01/meta_info autonomy_datasets_msgs/msg/ObjectListMetaInfo Annotations without a representation in perception_msgs/msg/Object: original_class. Associated with the object list via the header stamp and the object id.
Transformations /tf, /tf_static tf2_msgs/msg/TFMessage Static transformations to all sensor frames and dynamic transformation from map to vehicle frame.

Usage

Login using your HuggingFace Token to access the dataset and run the ROS node to download and store the data to rosbags while visualizing it in Rviz.

hf auth login
ros2 launch autonomy_datasets autonomy_datasets.launch.py dataset:=nvidia_physicalai_av_dataset

DrivIng Dataset

CC BY-NC-ND Harvard Dataverse

Rviz Screenshot DrivIng Dataset

DrivIng is a multimodal driving dataset recorded in Ingolstadt, Germany. The native data comprises the day, dusk, and night sequences, each synchronized at 10 Hz with a middle lidar, six vehicle cameras, vehicle state, calibration, and 3D track annotations. The dataset is licensed under CC BY-NC-ND 4.0.

Split Sequences
all night, day, dusk
day day
dusk dusk
night night
Source Topic Type Description
Sensor: Middle Lidar /lidar_01/point_cloud sensor_msgs/msg/PointCloud2 Point cloud in the middle-lidar frame with float32 fields (x, y, z, intensity) and a float64 absolute-seconds timestamp, preserving native per-point timing.
Sensor: Six Vehicle Cameras /camera_01/image_raw ... /camera_06/image_raw/camera_01/camera_info ... /camera_06/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo RGB images and native calibration; topic order is front-left, front-right, left, right, back-left, and back-right.
EgoData /ego_data perception_msgs/msg/EgoData Ego-vehicle pose in a local ENU map frame, derived from native relative north/east positions, north-referenced yaw, and the calibrated ADMA-to-vehicle lever arm. Velocity and standstill flag are differentiated from consecutive poses, as the native vehicle state holds no velocity.
Annotation: 3D Lidar Objects /object_list/lidar_01 perception_msgs/msg/ObjectList Track annotations as 3D objects in the middle-lidar frame.
Meta Information: Object Annotations /object_list/lidar_01/meta_info autonomy_datasets_msgs/msg/ObjectListMetaInfo Annotations without a representation in perception_msgs/msg/Object: original_class. Associated with the object list via the header stamp and the object id.
Transformations /tf, /tf_static tf2_msgs/msg/TFMessage Dynamic map to base_link pose plus calibrated static transforms to all sensors.

Usage

Select day, dusk, night, or all using dataset_split in params_driving.yml. Missing data is downloaded automatically from Harvard Dataverse and stored using the following folder structure:

$DATASET_DIR/
driving/
day/
annotations.json
calibration.json
timesync_info.csv
middle_lidar/
front_left_camera/
...
dusk/
night/

Run the ROS node to download, convert, and store the data to rosbags while visualizing it in Rviz.

ros2 launch autonomy_datasets autonomy_datasets.launch.py dataset:=driving

TUM Traffic Dataset

CC BY-NC-ND TUM Traffic Dataset

Rviz Screenshot TUM Traffic Dataset

The TUM Traffic Dataset (TUMTraf) is recorded by roadside sensors mounted on the gantry bridges of the Providentia++ test field along the A9 motorway and the S110 intersection near Munich, Germany. It is an infrastructure dataset without an ego vehicle, so no /ego_data is published; the sensor station is published as a static base_link. The dataset is licensed under CC BY-NC-ND 4.0.

The dataset is released as one archive per release and subset. All releases share a common file layout but differ in their sensors, directory names and label formats, so the adapter discovers the recordings, sensors and frame timestamps from the file names instead of hard-coding each release:

Release Subsets Sensors Annotations
R00 TUMTraf A9 Highway (image subsets) r00_s00 ... r00_s02 4 A9 gantry cameras (s040, s050) 3D box corners projected into the image — no object list published, see below
R00 TUMTraf A9 Highway (lidar subsets) r00_s03, r00_s04 Roadside lidars Native pre-OpenLABEL 3D cuboids (yaw-only orientation, no persistent track IDs)
R01 TUMTraf A9 Highway Extended r01_s01 ... r01_s04 4 A9 gantry cameras (s040, s050) 3D box corners projected into the image — no object list published, see below
R02 TUMTraf Intersection r02_s01 ... r02_s04 2 S110 cameras, 2 S110 Ouster lidars OpenLABEL 3D cuboids with track IDs

‍Only R00 to R02 of the TUM Traffic Dataset, which contains R00 to R05, are supported.

‍No object list for the R00 image subsets and R01: These releases annotate a 3D box only as its 8 corners projected into the 2D image (box3d_projected), without releasing the 3D pose (position, dimensions, orientation) that produced the projection. The dataset adapter does not attempt to recover that 3D pose. These recordings still publish their raw camera images, calibration and transforms. Only recordings with real 3D cuboids (the R00 lidar subsets, native pre-OpenLABEL format, and R02, OpenLABEL) publish /object_list/lidar_01.

The R00 lidar subsets ship no calibration source at all, so base_link is aliased to their single lidar's frame with an identity transform rather than leaving /tf_static unresolved.

Sensors are mapped onto the canonical topics in a fixed order, so camera_01 and lidar_01 are the sensors the object list is annotated in. The example below lists the topics of the intersection subsets (R02):

Source Topic Type Description
Sensor: South Lidar (Ouster OS1-64) /lidar_01/point_cloud sensor_msgs/msg/PointCloud2 Point cloud in the sensor frame with float32 fields (x, y, z, intensity) and a float64 absolute-seconds timestamp, preserving the native per-point timing.
Sensor: North Lidar (Ouster OS1-64) /lidar_02/point_cloud sensor_msgs/msg/PointCloud2 Point cloud in the sensor frame, fields as above.
Sensor: South1 Camera (Basler 8mm) /camera_01/image_raw/camera_01/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo RGB images (height=1200px, width=1920px) with the native calibration.
Sensor: South2 Camera (Basler 8mm) /camera_02/image_raw/camera_02/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo RGB images (height=1200px, width=1920px) with the native calibration.
Annotation: 3D Lidar Objects /object_list/lidar_01 perception_msgs/msg/ObjectList Annotated 3D objects (HEXAMOTION model) in the lidar_01 frame, with the track UUID, the native class and the native attributes (occlusion level, body color, number of points) in meta_info.
Transformations /tf, /tf_static tf2_msgs/msg/TFMessage Static transformations from the sensor station (base_link) to all sensor frames, and the station's static pose in the map frame.

‍The object list is only geometrically accurate against lidar_01: On releases with more than one lidar (R02 and newer), the objects are annotated directly in the reference lidar's own frame (lidar_01) and republished as-is; they are not re-derived per sensor. Overlaying /object_list/lidar_01 onto /lidar_02/point_cloud (or any other non-reference lidar) will show a visible offset, because the dataset's own released extrinsic calibration between its lidars is imprecise (e.g. for s110_lidar_ouster_south/s110_lidar_ouster_north in R02), which is why the dev kit ships a dedicated src/registration/point_cloud_registration.py to refine this pairing via ICP. This adapter does not run that registration step. So, only lidar_01 is aligned with the published objects.

Some tracked objects visibly float above or sink into the ground: On R02 and newer, a track's z and height are often set once and held constant for its whole lifetime while only x/y/yaw keep updating — confirmed directly in the raw label files, where most multi-frame tracks in a sample recording had byte-identical z/height despite moving tens of meters. This can leave a track that started well aligned drifting out of alignment later (e.g. over a stretch with different road elevation), or leave it wrong for its entire length if the frozen value was never accurate to begin with (e.g. estimated from only a handful of lidar points at long range and never revisited, even once the object is later observed with far denser support). This is a property of the dataset's own annotations, not of this adapter: cuboid.val is passed through per frame unmodified.

The sensors are triggered independently and the dataset ships no synchronization table, so each sample is built from the frames closest in time to the reference sensor (lidar_01, or camera_01 for the camera-only releases). Frames without a match within tum_traffic_sync_tolerance_seconds are skipped. Long recordings are split into rosbag scenes of tum_traffic_rosbag_duration_seconds.

All calibration is read from the dataset itself: from the _calibration directory of a recording if it ships one (R00/R01), otherwise from the coordinate_systems and streams sections of its OpenLABEL label files (R02 and newer).

Usage

The dataset cannot be downloaded automatically. Register, accept the license, and download the archives of the releases you want to use. Place the downloaded ZIP archives in the dataset directory; they are extracted on the first run into a directory named after the archive:

$DATASET_DIR/
tum_traffic/
a9_dataset_r02_s04.zip # placed here manually, extracted on the first run
a9_dataset_r02_s04/
images/
s110_camera_basler_south1_8mm/
*.jpg
...
point_clouds/
s110_lidar_ouster_south/
*.pcd
...
labels_point_clouds/
s110_lidar_ouster_south/
*.json
...
a9_dataset_r00_s00/ # or extracted manually
_images/
_labels/
_calibration/

Select the recordings to publish using dataset_split in params_tum_traffic.yml: all publishes every recording found in the dataset directory, any other value selects the recordings whose path contains it, e.g. a release (r02), a subset (r02_s04) or a split directory of a release (train). Because the releases ship different sensors, prefer a release-specific split; recordings of a mixed split publish their missing sensors as empty messages.

Run the ROS node to convert and store the data to rosbags while visualizing it in Rviz.

ros2 launch autonomy_datasets autonomy_datasets.launch.py dataset:=tum_traffic

Zenseact Open Dataset

CC BY-SA Zenseact Open Dataset

Rviz Screenshot Zenseact Open Dataset

The Zenseact Open Dataset (ZOD) is a multimodal driving dataset recorded by Zenseact over two years in 14 European countries. Its sensor suite is a single forward-looking 8 MP fisheye camera, three roof-mounted Velodyne lidars (one VLS128 and two VLP16) merged into one point cloud per scan, and an OxTS RT3000 GNSS/IMU. The dataset is released under a permissive license (CC BY-SA 4.0), which allows both research and commercial use.

ZOD is published as three sub-datasets, which are selected together with the version and the split through dataset_split in the form <subset>_<version>_<split>:

Sub-dataset Content Annotations
frames 100.000 independent keyframes from all over Europe, each with one camera image, one second of surrounding lidar scans in either direction and GNSS/IMU data Fully annotated
sequences 1.473 clips of 20 seconds with the complete sensor suite at 10 Hz Keyframe (middle frame) only
drives 29 clips of a few minutes with the complete sensor suite at 10 Hz Not annotated
Split Scenes Samples
frames_full_<train\|val\|all> 100.000 frames 1 per frame
sequences_full_<train\|val\|all> 1.473 sequences of 20 seconds approx. 200 per sequence
drives_full_<train\|val\|all> 29 drives of a few minutes approx. 10 per second
frames_mini_all 12 frames (10 train, 2 validation) 12
sequences_mini_all 2 sequences (1 train, 1 validation) approx. 400
drives_mini_all 2 drives (1 train, 1 validation) approx. 4.700

ZOD calibrates its sensors against an ISO-8855 reference frame at the center of the rear axle at ground level, which is published as base_link.

Source Topic Type Description
Sensor: Roof Lidars (1x Velodyne VLS128, 2x Velodyne VLP16) /lidar_01/point_cloud sensor_msgs/msg/PointCloud2 Point cloud in the lidar frame (approx. 254.000 points per scan) with float32 fields (x, y, z, intensity), a float64 absolute-seconds timestamp preserving the native per-point timing, and the uint8 diode_index identifying the emitter, and therefore the lidar, a point was measured by. ZOD merges the returns of all three lidars into a single scan.
Sensor: Front Camera (8 MP fisheye, 120° HFOV) /camera_01/image_raw/camera_01/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Anonymized RGB images (height=2168px, width=3848px) with the native Kannala-Brandt calibration, published as the equidistant distortion model.
EgoData /ego_data perception_msgs/msg/EgoData Ego-vehicle pose in a local ENU map frame, plus the velocity, acceleration and yaw rate of the high-precision GNSS/IMU interpolated onto the sample's timestamp.
Annotation: 3D Lidar Objects /object_list/lidar_01 perception_msgs/msg/ObjectList Annotated 3D objects (HEXAMOTION model) in the lidar_01 frame they are annotated in.
Annotation: 3D Camera Objects /object_list/camera_01 perception_msgs/msg/ObjectList The same objects, transformed into the camera_01 frame.
Meta Information: Object Annotations /object_list/lidar_01/meta_info/object_list/camera_01/meta_info autonomy_datasets_msgs/msg/ObjectListMetaInfo Annotations without a representation in perception_msgs/msg/Object: original_class, original_subclass, annotation_uuid, unclear, and object_type, occlusion_level, with_rider, emergency, artificial and traffic_content_visible wherever ZOD annotates them. Associated with the object list via the header stamp and the object id.
Transformations /tf, /tf_static tf2_msgs/msg/TFMessage Static transformations from the ISO-8855 vehicle frame (base_link) to the sensor frames, and the dynamic pose of base_link in the map frame.
Note
Only keyframes are annotated: ZOD annotates one keyframe per frame and per sequence, so exactly the sample recorded closest to that keyframe publishes the object lists. Every other sample of a sequence is published without the object list topics rather than with an empty object list, and the drives, which are not annotated at all, publish no object list at any sample. Set dataset_split to a frames split to obtain annotated samples only.

Every frame is its own rosbag scene: Frames are independent recordings taken in different countries and months, so consecutive frames are seconds to weeks apart. Each one therefore becomes a scene of its own, rather than being grouped, which would put a rosbag on a timeline that jumps between its samples and break every consumer running on the published /clock. Sequences and drives are continuous recordings and are split into scenes of zod_rosbag_duration_seconds instead.

Objects annotated in 2D only are not published: Every ZOD object carries a 2D box in the camera image, but roughly 70% of them additionally carry a 3D cuboid. Objects without a released 3D cuboid cannot be expressed as a perception_msgs/msg/Object and are left out of the object lists.

Static roadside objects are published as UNKNOWN: ZOD annotates poles, traffic signs, traffic signals, traffic guides and dynamic barriers alongside vehicles and vulnerable road users. perception_msgs/msg/ObjectClassification has no class for them, so they are classified as UNKNOWN, i.e. "definitely none of the other defined classes"; their ZOD class is preserved in original_class and original_subclass. Objects that ZOD flags as unclear are published as UNCLASSIFIED.

The remaining annotation projects are not published: ZOD also ships lane marking and ego road segmentation, a traffic sign taxonomy of 156 classes, and road condition labels. These have no representation in perception_msgs and are not converted. Radar, which later ZOD releases add for sequences and drives, is not converted either.

The ego vehicle dimensions are an approximation: ZOD publishes no dimensions for its collection vehicles, so EgoData reports the dimensions of a large passenger estate car, consistent with the released calibration and with the ego-return box of the development kit.
For this dataset, Zenseact AB has taken all reasonable measures to remove all personally identifiable information, including faces and license plates. To the extent that you like to request removal of specific images from the dataset, please contact priva.nosp@m.cy@z.nosp@m.ensea.nosp@m.ct.c.nosp@m.om.

The camera runs at 10.1 Hz and the lidar at 9 Hz, and ZOD ships no synchronization table, so each sample is built from the frames closest in time to the reference sensor, which is the camera because ZOD defines the camera images as its keyframes. Frames without a match within zod_sync_tolerance_seconds are skipped, which typically drops the first sample of a sequence. Point clouds are motion-compensated onto the sample's timestamp, so that lidar, camera and annotations describe the same instant.

The poses ZOD publishes are relative to the first GNSS/IMU sample of a scene, with the x axis along the ego vehicle's heading at that sample. They are rotated by that heading, which is read from the scene's oxts.hdf5, so that map is an ENU frame anchored at that first sample. Scenes of the frames sub-dataset are independent recordings from different places, so their map frames are unrelated to each other.

Usage

The dataset requires registration: apply for access to receive a personal download link. Set it via the zod_download_url parameter in params_zenseact_open_dataset.yml or via the ZOD_DOWNLOAD_URL environment variable, and the node downloads and extracts the selected sub-dataset on the first run. Alternatively, download it manually with the CLI of the development kit:

zod download -y --url="<download-link>" --output-dir=$DATASET_DIR/zenseact_open_dataset --subset=frames --version=mini

Both ways produce the following folder structure, in which all three sub-datasets live next to each other:

$DATASET_DIR/
zenseact_open_dataset/
trainval-frames-mini.json
single_frames/
044953/
calibration.json
ego_motion.json
metadata.json
oxts.hdf5
annotations/
*.json
camera_front_blur/
*.jpg
camera_front_dnat/
*.jpg
lidar_velodyne/
*.npy
...
trainval-sequences-mini.json
sequences/
000002/
...
trainval-drives-mini.json
drives/
000005/
...

‍Sub-datasets downloaded into separate directories are picked up as well: an index that is not found in the dataset directory itself is also looked up one level below it, e.g. $DATASET_DIR/zenseact_open_dataset/frames_mini/trainval-frames-mini.json.

Select the sub-dataset, version and split using dataset_split in params_zenseact_open_dataset.yml, e.g. frames_mini_val, sequences_full_train or drives_mini_all. ZOD provides two anonymizations of its camera images, deep fake anonymization (dnat) and blurring (blur); the frames sub-dataset ships both and is selected via zod_anonymization, while sequences and drives ship the blurred images only.

Run the ROS node to download, convert, and store the data to rosbags while visualizing it in Rviz.

ros2 launch autonomy_datasets autonomy_datasets.launch.py dataset:=zenseact_open_dataset

FZI-AURA Dataset

CC BY-SA FZI-AURA

Rviz Screenshot FZI-AURA Dataset

FZI-AURA is a multimodal driving dataset recorded across southern Germany with CoCar NextGen, the research vehicle of the FZI Research Center for Information Technology. It carries the largest lidar suite of any public automated driving dataset: six rotating Ouster lidars (4x OS1-64, 2x OS2-128) and six Aeva Aeries II FMCW lidars give 360° coverage twice over, complemented by eight global-shutter surround-view cameras, up to three Continental ARS 548 RDI radars and an INS. Besides 3D boxes it ships more than 30 billion human-annotated semantic lidar points, more than any other public non-synthetic driving dataset. It is released under a permissive license (CC BY-SA 4.0), which allows both research and commercial use.

Split Scenes Samples
train 1.979 approx. 200 per scene
val 247 approx. 200 per scene
test 247 approx. 200 per scene
all 2.473 493.754

Each scene is a self-contained recording of roughly 20 seconds sampled at 10 Hz and becomes one rosbag scene. FZI-AURA annotates 2 Hz keyframes only and releases the sensor payloads of keyframes and non-keyframes as separate download layers, which fzi_aura_samples selects between: keyframes publishes the annotated 2 Hz samples covered by the default download, all publishes the full 10 Hz sample stream.

The sensor suite differs between scenes (4 to 12 lidars, 0 to 3 radars), so a sensor is only mapped to a topic when at least one selected scene holds it. The canonical topic of a sensor is fixed by its position in the suite, so the same sensor always reaches the same topic:

Topic Camera Topic Lidar Topic Radar
camera_01 front_medium lidar_01 top_left radar_01 front_left
camera_02 front_wide lidar_02 top_right radar_02 front_right
camera_03 front_tele lidar_03 front_left radar_03 rear_center
camera_04 right_forward lidar_04 front_right
camera_05 right_rearward lidar_05 rear_left
camera_06 rear_wide lidar_06 rear_right
camera_07 left_rearward lidar_07 aeva_front_center
camera_08 left_forward lidar_08 aeva_front_left
lidar_09 aeva_front_right
lidar_10 aeva_side_left
lidar_11 aeva_side_right
lidar_12 aeva_rear_center

FZI-AURA calibrates its sensors against base_link, which already follows the ROS convention (x forward, y left, z up) and is published unchanged. Sensor frames are published as <modality>_<sensor id>, e.g. lidar_top_left or radar_front_left, because a bare sensor ID is not unique across modalities.

Source Topic Type Description
Sensor: Ouster Lidars (4x OS1-64, 2x OS2-128) /lidar_01/point_cloud.../lidar_06/point_cloud sensor_msgs/msg/PointCloud2 Point cloud in the lidar frame (approx. 60.000 points per scan) with the native fields (x, y, z, t, reflectivity, ring, range). Motion-compensated onto the end of the sweep by default, selectable via fzi_aura_lidar_stage.
Sensor: Aeva Aeries II FMCW Lidars /lidar_07/point_cloud.../lidar_12/point_cloud sensor_msgs/msg/PointCloud2 Point cloud in the lidar frame with the native fields of the FMCW sensor, including the per-point Doppler velocity measured along the beam and an intensity next to the reflectivity. The raw and the motion-compensated stage carry different fields.
Sensor: Surround-View Cameras (8x global shutter) /camera_01/image_raw.../camera_08/image_raw/camera_0X/camera_info sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo Anonymized RGB images, 1920x1200px for the six narrow cameras and 2592x2048px for front_wide and rear_wide, scalable via fzi_aura_image_scale. The images are rectified and are published with their projection matrix and zero distortion coefficients.
Sensor: Radars (up to 3x Continental ARS 548 RDI) /radar_01/point_cloud.../radar_03/point_cloud sensor_msgs/msg/PointCloud2 Radar detections in the radar frame with the fields (x, y, z, rcs, elevation_angle) and, where the sensor reports it, the radial velocity range_rate.
Annotation: Semantic Lidar Labels /lidar_01/point_cloud.../lidar_06/point_cloud sensor_msgs/msg/PointCloud2 Per-point semantic_id and instance_id fields of the cloud they annotate, added wherever the scene is semantically labeled. Controlled via fzi_aura_publish_semantic_labels.
EgoData /ego_data perception_msgs/msg/EgoData Ego-vehicle dimensions and dynamics state (EGO model) in an ENU map frame. Pose from the dataset's ego poses, velocity, acceleration, yaw rate, steering angle, standstill flag and turn indicator from the INS and the vehicle's CAN bus.
Annotation: 3D Lidar Objects /object_list/lidar_01 perception_msgs/msg/ObjectList Annotated 3D objects (HEXAMOTION model) in the frame of the reference lidar, restricted to the objects holding at least one point of that lidar.
Annotation: 3D Vehicle-Frame Objects /object_list/base_link perception_msgs/msg/ObjectList The canonical annotation: every 3D object of the keyframe in the base_link frame, including objects no single lidar sees points of.
Meta Information: Object Annotations /object_list/lidar_01/meta_info/object_list/base_link/meta_info autonomy_datasets_msgs/msg/ObjectListMetaInfo Annotations without a representation in perception_msgs/msg/Object: original_class, the object_id the dataset tracks an object under within a scene, and the sensor_id of the per-sensor object list. Associated with the object list via the header stamp and the object id.
Transformations /tf, /tf_static tf2_msgs/msg/TFMessage Static transformations from the vehicle frame (base_link) to every calibrated sensor frame, and the dynamic pose of base_link in the map frame.
Map map_contents (parameter) string Lanelet2 map (OSM XML) of the roads around the current scene, generated from OpenStreetMap. Updated on every scene change, analogous to lanelet2_map_server.

FZI-AURA ships no map, so the Lanelet2 map of a scene is generated from the OpenStreetMap roads within 200 m of its GNSS track, which are fetched from the Overpass API configured via fzi_aura_overpass_url. The map generation requires internet access and is controlled via the publish_lanelet2_map (enable/disable) and fzi_aura_lanelet2_lane_width (assumed lane width in meters) parameters. A scene that cannot be fetched is published without a map. The map is stored next to the rosbag data of its scene like the nuScenes map, so replaying a rosbag does not fetch it again.

Each road becomes one road lanelet per lane (highway for motorways, play_street for living streets), laid out around its OSM centerline from the lanes, lanes:forward, lanes:backward and oneway tags for right-hand traffic; the lane boundaries are synthesized by offsetting the centerline by multiples of the lane width. Crossings mapped as footway=crossing become crosswalk lanelets. The map is georeferenced into the map frame by the GNSS fixes of the scene's samples, which agree with the ego poses to a few decimeters.

Note
The map approximates the road layout: OpenStreetMap maps roads rather than lanes. Lane widths are assumed, lanes that are not tagged are guessed (one per direction on two-way roads), and the lanelets of different roads end at their shared junction node rather than being connected through the junction. OpenStreetMap geometry is typically accurate to one or two meters. The generated maps contain information from OpenStreetMap, © OpenStreetMap contributors, available under the Open Database License.

Only keyframes are annotated: 3D boxes and semantic labels are provided at 2 Hz keyframes, while the sample stream runs at 10 Hz. A non-keyframe sample is published without the object list topics rather than with an empty object list; an annotated keyframe holding no object does publish an empty one, which states that nothing was annotated in it. fzi_aura_samples: keyframes publishes the annotated samples only.

Non-keyframe samples need their own download layers: The default download covers the keyframe payloads. With fzi_aura_samples: all, the samples between two keyframes are published without sensor data unless the camera_nonkeyframes, lidar_raw_nonkeyframes and radar_nonkeyframes layers were downloaded as well. Motion-compensated lidar is not released for non-keyframes at all.

The sensor suite differs between scenes: Most scenes hold the six Ouster lidars, and a subset additionally holds the six Aeva lidars. A scene that does not hold a sensor is published without its topics; the topics stay advertised as long as any selected scene holds the sensor.

The map frame is aligned with the INS attitude: The frame the released ego poses are expressed in carries an arbitrary orientation per recording that is neither gravity-aligned nor north-referenced — the poses of a scene can hold a roll of more than ten degrees while the vehicle drives level. The INS state in the vehicle signals is a proper east-north-up attitude, so the poses are rotated by the offset between the two at the first sample of a scene, which yields an ENU-aligned map. The residual drift of the released odometry stays below about two degrees over a scene. Scenes without vehicle signals keep the native orientation of the released poses.

Riders and micromobility share a class with their vehicle: FZI-AURA annotates two-wheelers both as the bare vehicle (bicycle, motorcycle, portable) and as the vehicle together with its rider (bicyclist, motorcyclist, portable-rider). perception_msgs/msg/ObjectClassification defines BICYCLE, MOTORCYCLE and MICRO as covering the vehicle and its rider, so both spellings map to the same class; a rider box holds the person alone and is published as VRU. Objects of the dynamic class, which collects movable objects that fit none of the other classes, are published as UNKNOWN. The dataset's own class is preserved in original_class.

Usage

The dataset is available on Hugging Face: after log in, the node will download the selected scenes on the first run via the FZI-AURA SDK downloader.

hf auth login

Alternatively, download the data manually with the CLI of the SDK:

fzi-aura-download $DATASET_DIR/fzi_aura --splits val --scenes "2025-06-13-07-09-37|75"

Both ways produce the following folder structure:

$DATASET_DIR/
fzi_aura/
dataset.json
available_data.json
splits/
v1.0/
train.txt
val.txt
test.txt
scenes/
2025-06-13-07-09-37_75/
scene.json
samples.jsonl
samples.parquet
calibration.json
camera/
front_medium/
*.jpg
...
lidar/
raw/
top_left/
*.pcd
...
motion_compensated/
...
radar/
front_left/
*.pcd
...
labels/
boxes_3d.jsonl
boxes_3d_sensor_frame/
top_left/
boxes_3d.jsonl
...
semantic/
top_left/
*.label
...
classes.json
ego/
poses.parquet
vehicle_signals.parquet
...

Select the split and the scenes using dataset_split and fzi_aura_scenes in params_fzi_aura.yml.

Run the ROS node to download, convert, and store the data to rosbags while visualizing it in Rviz.

ros2 launch autonomy_datasets autonomy_datasets.launch.py dataset:=fzi_aura

Thinking Cars Dataset

commercial Thinking Cars

Custom datasets according to your needs and suitable for commercial use are available via an expanding network of partners on request, for example:

  • Sensor data from (stereo) cameras, lidars, radars and IMU
  • Object annotations
  • V2X Data (e.g. ETSI ITS Messages)
  • Driving Trajectories and Scenarios

Adding a new dataset

  1. Create a new dataset adapter based on the existing files here.
  2. Publish annotations that have no representation in perception_msgs/msg/Object (e.g. the dataset's original class name) as autonomy_datasets_msgs/msg/ObjectListMetaInfo on the object list's meta_info topic, using the helpers in meta_info.py.
  3. Add documentation for the new dataset to this README and add it to the table in the top-level README.
  4. Create a Pull Request on GitHub and wait for maintainer's feedback.