|
autonomy_datasets v1.6.0
|
This repository supports various automated driving datasets including:

| Split | Samples |
|---|---|
all | 198.068 |
training | 158.081 |
validation | 39.987 |
| Source | Topic | Type | Description |
|---|---|---|---|
| Sensor: Top Lidar | /lidar_01/point_cloud | sensor_msgs/msg/PointCloud2 | Raw sensor data from top lidar as point cloud with fields (x, y, z, intensity, elongation) in sensor frame. |
| Sensor: Front Camera | /camera_01/image_raw/camera_01/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=1280px, width=1920px) from front camera. |
| Sensor: Front-Left Camera | /camera_02/image_raw/camera_02/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=1280px, width=1920px) from front-left camera. |
| Sensor: Front-Right Camera | /camera_03/image_raw/camera_03/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=1280px, width=1920px) from front-right camera. |
| Sensor: Side-Left Camera | /camera_04/image_raw/camera_04/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=886px, width=1920px) from side-left camera. |
| Sensor: Side-Right Camera | /camera_05/image_raw/camera_05/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=886px, width=1920px) from side-right camera. |
| EgoData | /ego_data | perception_msgs/msg/EgoData | Ego-vehicle's dimensions and dynamics state in map frame. |
| Annotation: 3D Lidar Objects | /object_list/lidar_01 | perception_msgs/msg/ObjectList | Annotated 3D objects (HEXAMOTION model) in vehicle frame. Default: Only objects with min. 1 point in top lidar point cloud. |
| Annotation: 2D Camera Objects | /object_list/cameras | perception_msgs/msg/ObjectList | Annotated 2D objects (CAMERA2D model) in camera frame. Note: Currently no visualization is shown for this data type in RViz. |
| Meta Information: Object Annotations | /object_list/lidar_01/meta_info/object_list/camera_01/meta_info/object_list/camera_all/meta_info | autonomy_datasets_msgs/msg/ObjectListMetaInfo | Annotations without a representation in perception_msgs/msg/Object: original_class, num_lidar_pts and difficulty_level. Associated with the object list via the header stamp and the object id. |
| Transformations | /tf, /tf_static | tf2_msgs/msg/TFMessage | Static transformations to all sensor frames and dynamic transformation from map to vehicle frame. |
Download the dataset and ensure the following folder structure is correct:
Run the ROS node to convert and store the data to rosbags while visualizing it in Rviz.

| Split | Samples |
|---|---|
training | 28.130 |
validation | 6.019 |
training_mini | 25 |
validation_mini | 17 |
| Source | Topic | Type | Description |
|---|---|---|---|
| Sensor: Top Lidar (Velodyne HDL-32E) | /lidar_01/point_cloud | sensor_msgs/msg/PointCloud2 | Raw sensor data from top lidar as point cloud with fields (x, y, z, intensity, timestamp). |
| Sensor: Front Camera (Basler acA1600-60gc) | /camera_01/image_raw/camera_01/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=900px, width=1600px) from front camera. |
| Sensor: Front-Right Camera (Basler acA1600-60gc) | /camera_02/image_raw/camera_02/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=900px, width=1600px) from front-right camera. |
| Sensor: Back-Right Camera (Basler acA1600-60gc) | /camera_03/image_raw/camera_03/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=900px, width=1600px) from back-right camera. |
| Sensor: Back Camera (Basler acA1600-60gc) | /camera_04/image_raw/camera_04/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=900px, width=1600px) from back camera. |
| Sensor: Back-Left Camera (Basler acA1600-60gc) | /camera_05/image_raw/camera_05/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=900px, width=1600px) from back-left camera. |
| Sensor: Front-Left Camera (Basler acA1600-60gc) | /camera_06/image_raw/camera_06/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=900px, width=1600px) from front-left camera. |
| Sensor: Front Radar (Continental ARS 408-21) | /radar_01/point_cloud | sensor_msgs/msg/PointCloud2 | Radar detections from front radar as point cloud with fields (x, y, z, radial_velocity, rcs). |
| Sensor: Front-Right Radar (Continental ARS 408-21) | /radar_02/point_cloud | sensor_msgs/msg/PointCloud2 | Radar detections from front-right radar as point cloud with fields (x, y, z, radial_velocity, rcs). |
| Sensor: Back-Right Radar (Continental ARS 408-21) | /radar_03/point_cloud | sensor_msgs/msg/PointCloud2 | Radar detections from back-right radar as point cloud with fields (x, y, z, radial_velocity, rcs). |
| Sensor: Back-Left Radar (Continental ARS 408-21) | /radar_04/point_cloud | sensor_msgs/msg/PointCloud2 | Radar detections from back-left radar as point cloud with fields (x, y, z, radial_velocity, rcs). |
| Sensor: Front-Left Radar (Continental ARS 408-21) | /radar_05/point_cloud | sensor_msgs/msg/PointCloud2 | Radar detections from front-left radar as point cloud with fields (x, y, z, radial_velocity, rcs). |
| EgoData | /ego_data | perception_msgs/msg/EgoData | Ego-vehicle's dimensions and dynamics state (EGO model) in map frame. Pose from the dataset's ego poses, velocity, acceleration, yaw rate, steering angle, standstill flag, turn indicator and brake light from the CAN bus expansion. |
| Annotation: 3D Lidar Objects | /object_list/lidar_01 | perception_msgs/msg/ObjectList | Annotated 3D objects (HEXAMOTION model) visible in lidar scan. As nuScenes stores no object dynamics, the absolute velocity, acceleration and yaw rate are differentiated from the annotation positions and yaw angles of the neighboring keyframes. |
| Annotation: 3D Front Camera Objects | /object_list/camera_01 | perception_msgs/msg/ObjectList | Annotated 3D objects (HEXAMOTION model) visible in front camera image. As nuScenes stores no object dynamics, the absolute velocity, acceleration and yaw rate are differentiated from the annotation positions and yaw angles of the neighboring keyframes. |
| Transformations | /tf, /tf_static | tf2_msgs/msg/TFMessage | Static transformations to all sensor frames and dynamic transformation from map to vehicle frame. |
| Detection: Detected 3D Objects | /object_list/detected | perception_msgs/msg/ObjectList | Detected 3D objects (HEXAMOTION model) from Megvii baseline. |
| Meta Information: Object Annotations | /object_list/lidar_01/meta_info/object_list/camera_01/meta_info/object_list/detected/meta_info | autonomy_datasets_msgs/msg/ObjectListMetaInfo | Annotations without a representation in perception_msgs/msg/Object: original_class, num_lidar_pts, num_radar_pts, num_points, attribute and detection_score. Associated with the object list via the header stamp and the object id. |
| Map | map_contents (parameter) | string | Lanelet2 map (OSM XML) of the current scene's location, converted from the nuScenes map expansion. Updated on every scene change, analogous to lanelet2_map_server. |
The Lanelet2 conversion is controlled via the publish_lanelet2_map (enable/disable) and nuscenes_lanelet2_lane_width (assumed lane width in meters) parameters. Lanes and lane connectors are converted to road lanelets (boundaries synthesized by offsetting the centerline by half the lane width), and pedestrian crossings to crosswalk lanelets. Conversion requires the nuScenes map-expansion data under maps/expansion/.
The converted map is stored next to the rosbag data of the scene it belongs to: the map as map.osm and its origin as map.yaml inside the scene's rosbag directory. Replaying a rosbag restores the map from there instead of converting it again, unless publish_lanelet2_map is disabled.
Download the dataset (including CAN Bus and Map Expansion) and ensure the following folder structure is correct:
Run the ROS node to convert and store the data to rosbags while visualizing it in Rviz.

MAN TruckScenes is a public dataset recorded from a heavy truck. It comprises 747 scenes of 20 seconds each, annotated at 2 Hz, recorded with 6 lidars, 6 radars, 4 cameras and a high-precision GNSS. The dataset reuses the nuScenes database schema and is licensed under CC BY-NC-SA 4.0.
| Split | Scenes | Samples |
|---|---|---|
train | 523 | approx. 20.900 |
val | 75 | approx. 3.000 |
test | 149 | approx. 5.900 |
mini_train | 8 | approx. 320 |
mini_val | 2 | approx. 80 |
The
testsplit is released without object annotations, so the object list topics are published empty for it.
| Source | Topic | Type | Description |
|---|---|---|---|
| Sensor: 6 Lidars | /lidar_01/point_cloud ... /lidar_06/point_cloud | sensor_msgs/msg/PointCloud2 | Point clouds in the respective sensor frame with float32 fields (x, y, z, intensity) and a float64 absolute-seconds timestamp, preserving native per-point timing. Topic order is left, right, top-front, top-left, top-right, and rear. |
| Sensor: 6 Radars | /radar_01/point_cloud ... /radar_06/point_cloud | sensor_msgs/msg/PointCloud2 | Radar detections in the respective sensor frame with float32 fields (x, y, z, vrel_x, vrel_y, vrel_z, rcs). Topic order is left-front, right-front, right-side, right-back, left-back, and left-side. |
| Sensor: 4 Cameras | /camera_01/image_raw ... /camera_04/image_raw/camera_01/camera_info ... /camera_04/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Undistorted and rectified RGB images with native calibration; topic order is left-front, right-front, right-back, and left-back. |
| EgoData | /ego_data | perception_msgs/msg/EgoData | Ego-vehicle's dimensions and dynamics state in the UTM-WGS84 (zone U32) map frame, with velocities, accelerations, and yaw rate taken from the native ego_motion_chassis table. |
| Annotation: 3D Lidar Objects | /object_list/lidar_01 | perception_msgs/msg/ObjectList | Annotated 3D objects (HEXAMOTION model) in the left lidar frame. Default: Only objects with min. 1 point in the lidar point cloud. |
| Annotation: 3D Camera Objects | /object_list/camera_01 | perception_msgs/msg/ObjectList | Annotated 3D objects (HEXAMOTION model) visible in the left front camera image. |
| Meta Information: Object Annotations | /object_list/lidar_01/meta_info/object_list/camera_01/meta_info | autonomy_datasets_msgs/msg/ObjectListMetaInfo | Annotations without a representation in perception_msgs/msg/Object: original_class, num_lidar_pts, num_radar_pts, num_points and attribute. Associated with the object list via the header stamp and the object id. |
| Transformations | /tf, /tf_static | tf2_msgs/msg/TFMessage | Static transformations to all sensor frames and dynamic transformation from map to vehicle frame. |
Select the split using dataset_split in params_truckscenes.yml. Missing data is downloaded automatically from the AWS Open Data registry without requiring credentials; alternatively download the archives manually and unpack them into the following folder structure:
Only keyframe sensor data (
samples/) is downloaded, because the adapter publishes annotated keyframes only. The unannotatedsweeps/are skipped to keep the required disk space low.
Run the ROS node to download, convert, and store the data to rosbags while visualizing it in Rviz.

The number of samples depends on the configurable selected sensor modalities:
| Sensor Modalities | Sensor Setup | Samples | | ---— | ---— | -— | | Camera | 7 cameras at 30 Hz | 306.152 (20 seconds each) | 183.691.200 | | Camera + Lidar | 7 cameras + 360 deg lidar at 10 Hz | 298.326 (20 seconds each) | 59.665.200 | | Camera + Radar | 7 camera + up to 10 radars at 10 Hz | 160.761 (20 seconds each) | 32.152.200 | | Camera + Lidar + Radar | 7 camera + 360 deg lidar at 10 Hz + up to 10 radars at 10 Hz | TODO (20 seconds each) | TODO |
The provided default splits contain only samples including all sensor modalities (Camera + Lidar + Radar).
| Split | Country | Scenes | Samples |
|---|---|---|---|
all | All | 85.082 | approx. 17.016.400 |
all | Germany | 7.247 | approx. 1.449.400 |
train | Germany | 3.694 | approx. 738.800 |
valid | Germany | 2.044 | approx. 408.800 |
test | Germany | 1.509 | approx. 301.800 |
| Source | Topic | Type | Description |
|---|---|---|---|
| Sensor: Top Lidar | /lidar_01/point_cloud | sensor_msgs/msg/PointCloud2 | Raw sensor data from top lidar as point cloud with fields (x, y, z, intensity) in sensor frame. |
| Sensor: Front Tele Camera (30° FOV) | /camera_01/image_raw/camera_01/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=1080px, width=1920px) from front tele camera. |
| Sensor: Front Wide Camera (120° FOV) | /camera_02/image_raw/camera_02/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=1080px, width=1920px) from front wide camera. |
| Sensor: Left Cross Camera (120° FOV) | /camera_03/image_raw/camera_03/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=1080px, width=1920px) from left cross camera. |
| Sensor: Right Cross Camera (120° FOV) | /camera_04/image_raw/camera_04/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=1080px, width=1920px) from right cross camera. |
| Sensor: Rear-Left Camera (70° FOV) | /camera_05/image_raw/camera_05/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=1080px, width=1920px) from rear-left camera. |
| Sensor: Rear-Right Camera (70° FOV) | /camera_06/image_raw/camera_06/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=1080px, width=1920px) from rear-right camera. |
| Sensor: Rear Tele Camera (30° FOV) | /camera_07/image_raw/camera_07/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Raw RGB images (height=1080px, width=1920px) from rear tele camera. |
| EgoData | /ego_data | perception_msgs/msg/EgoData | Ego-vehicle's dimensions and dynamics state in map frame. |
| Annotation: 3D Lidar Objects | /object_list/lidar_01 | perception_msgs/msg/ObjectList | Annotated 3D objects (HEXAMOTION model) in vehicle frame. |
| Meta Information: Object Annotations | /object_list/lidar_01/meta_info | autonomy_datasets_msgs/msg/ObjectListMetaInfo | Annotations without a representation in perception_msgs/msg/Object: original_class. Associated with the object list via the header stamp and the object id. |
| Transformations | /tf, /tf_static | tf2_msgs/msg/TFMessage | Static transformations to all sensor frames and dynamic transformation from map to vehicle frame. |
Login using your HuggingFace Token to access the dataset and run the ROS node to download and store the data to rosbags while visualizing it in Rviz.

DrivIng is a multimodal driving dataset recorded in Ingolstadt, Germany. The native data comprises the day, dusk, and night sequences, each synchronized at 10 Hz with a middle lidar, six vehicle cameras, vehicle state, calibration, and 3D track annotations. The dataset is licensed under CC BY-NC-ND 4.0.
| Split | Sequences |
|---|---|
all | night, day, dusk |
day | day |
dusk | dusk |
night | night |
| Source | Topic | Type | Description |
|---|---|---|---|
| Sensor: Middle Lidar | /lidar_01/point_cloud | sensor_msgs/msg/PointCloud2 | Point cloud in the middle-lidar frame with float32 fields (x, y, z, intensity) and a float64 absolute-seconds timestamp, preserving native per-point timing. |
| Sensor: Six Vehicle Cameras | /camera_01/image_raw ... /camera_06/image_raw/camera_01/camera_info ... /camera_06/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | RGB images and native calibration; topic order is front-left, front-right, left, right, back-left, and back-right. |
| EgoData | /ego_data | perception_msgs/msg/EgoData | Ego-vehicle pose in a local ENU map frame, derived from native relative north/east positions, north-referenced yaw, and the calibrated ADMA-to-vehicle lever arm. Velocity and standstill flag are differentiated from consecutive poses, as the native vehicle state holds no velocity. |
| Annotation: 3D Lidar Objects | /object_list/lidar_01 | perception_msgs/msg/ObjectList | Track annotations as 3D objects in the middle-lidar frame. |
| Meta Information: Object Annotations | /object_list/lidar_01/meta_info | autonomy_datasets_msgs/msg/ObjectListMetaInfo | Annotations without a representation in perception_msgs/msg/Object: original_class. Associated with the object list via the header stamp and the object id. |
| Transformations | /tf, /tf_static | tf2_msgs/msg/TFMessage | Dynamic map to base_link pose plus calibrated static transforms to all sensors. |
Select day, dusk, night, or all using dataset_split in params_driving.yml. Missing data is downloaded automatically from Harvard Dataverse and stored using the following folder structure:
Run the ROS node to download, convert, and store the data to rosbags while visualizing it in Rviz.

The TUM Traffic Dataset (TUMTraf) is recorded by roadside sensors mounted on the gantry bridges of the Providentia++ test field along the A9 motorway and the S110 intersection near Munich, Germany. It is an infrastructure dataset without an ego vehicle, so no /ego_data is published; the sensor station is published as a static base_link. The dataset is licensed under CC BY-NC-ND 4.0.
The dataset is released as one archive per release and subset. All releases share a common file layout but differ in their sensors, directory names and label formats, so the adapter discovers the recordings, sensors and frame timestamps from the file names instead of hard-coding each release:
| Release | Subsets | Sensors | Annotations |
|---|---|---|---|
R00 TUMTraf A9 Highway (image subsets) | r00_s00 ... r00_s02 | 4 A9 gantry cameras (s040, s050) | 3D box corners projected into the image — no object list published, see below |
R00 TUMTraf A9 Highway (lidar subsets) | r00_s03, r00_s04 | Roadside lidars | Native pre-OpenLABEL 3D cuboids (yaw-only orientation, no persistent track IDs) |
R01 TUMTraf A9 Highway Extended | r01_s01 ... r01_s04 | 4 A9 gantry cameras (s040, s050) | 3D box corners projected into the image — no object list published, see below |
R02 TUMTraf Intersection | r02_s01 ... r02_s04 | 2 S110 cameras, 2 S110 Ouster lidars | OpenLABEL 3D cuboids with track IDs |
Only
R00toR02of the TUM Traffic Dataset, which containsR00toR05, are supported.
No object list for the
R00image subsets andR01: These releases annotate a 3D box only as its 8 corners projected into the 2D image (box3d_projected), without releasing the 3D pose (position, dimensions, orientation) that produced the projection. The dataset adapter does not attempt to recover that 3D pose. These recordings still publish their raw camera images, calibration and transforms. Only recordings with real 3D cuboids (theR00lidar subsets, native pre-OpenLABEL format, andR02, OpenLABEL) publish/object_list/lidar_01.The
R00lidar subsets ship no calibration source at all, sobase_linkis aliased to their single lidar's frame with an identity transform rather than leaving/tf_staticunresolved.
Sensors are mapped onto the canonical topics in a fixed order, so camera_01 and lidar_01 are the sensors the object list is annotated in. The example below lists the topics of the intersection subsets (R02):
| Source | Topic | Type | Description |
|---|---|---|---|
| Sensor: South Lidar (Ouster OS1-64) | /lidar_01/point_cloud | sensor_msgs/msg/PointCloud2 | Point cloud in the sensor frame with float32 fields (x, y, z, intensity) and a float64 absolute-seconds timestamp, preserving the native per-point timing. |
| Sensor: North Lidar (Ouster OS1-64) | /lidar_02/point_cloud | sensor_msgs/msg/PointCloud2 | Point cloud in the sensor frame, fields as above. |
| Sensor: South1 Camera (Basler 8mm) | /camera_01/image_raw/camera_01/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | RGB images (height=1200px, width=1920px) with the native calibration. |
| Sensor: South2 Camera (Basler 8mm) | /camera_02/image_raw/camera_02/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | RGB images (height=1200px, width=1920px) with the native calibration. |
| Annotation: 3D Lidar Objects | /object_list/lidar_01 | perception_msgs/msg/ObjectList | Annotated 3D objects (HEXAMOTION model) in the lidar_01 frame, with the track UUID, the native class and the native attributes (occlusion level, body color, number of points) in meta_info. |
| Transformations | /tf, /tf_static | tf2_msgs/msg/TFMessage | Static transformations from the sensor station (base_link) to all sensor frames, and the station's static pose in the map frame. |
The object list is only geometrically accurate against
lidar_01: On releases with more than one lidar (R02and newer), the objects are annotated directly in the reference lidar's own frame (lidar_01) and republished as-is; they are not re-derived per sensor. Overlaying/object_list/lidar_01onto/lidar_02/point_cloud(or any other non-reference lidar) will show a visible offset, because the dataset's own released extrinsic calibration between its lidars is imprecise (e.g. fors110_lidar_ouster_south/s110_lidar_ouster_northinR02), which is why the dev kit ships a dedicatedsrc/registration/point_cloud_registration.pyto refine this pairing via ICP. This adapter does not run that registration step. So, onlylidar_01is aligned with the published objects.Some tracked objects visibly float above or sink into the ground: On
R02and newer, a track'szandheightare often set once and held constant for its whole lifetime while onlyx/y/yawkeep updating — confirmed directly in the raw label files, where most multi-frame tracks in a sample recording had byte-identicalz/heightdespite moving tens of meters. This can leave a track that started well aligned drifting out of alignment later (e.g. over a stretch with different road elevation), or leave it wrong for its entire length if the frozen value was never accurate to begin with (e.g. estimated from only a handful of lidar points at long range and never revisited, even once the object is later observed with far denser support). This is a property of the dataset's own annotations, not of this adapter:cuboid.valis passed through per frame unmodified.
The sensors are triggered independently and the dataset ships no synchronization table, so each sample is built from the frames closest in time to the reference sensor (lidar_01, or camera_01 for the camera-only releases). Frames without a match within tum_traffic_sync_tolerance_seconds are skipped. Long recordings are split into rosbag scenes of tum_traffic_rosbag_duration_seconds.
All calibration is read from the dataset itself: from the _calibration directory of a recording if it ships one (R00/R01), otherwise from the coordinate_systems and streams sections of its OpenLABEL label files (R02 and newer).
The dataset cannot be downloaded automatically. Register, accept the license, and download the archives of the releases you want to use. Place the downloaded ZIP archives in the dataset directory; they are extracted on the first run into a directory named after the archive:
Select the recordings to publish using dataset_split in params_tum_traffic.yml: all publishes every recording found in the dataset directory, any other value selects the recordings whose path contains it, e.g. a release (r02), a subset (r02_s04) or a split directory of a release (train). Because the releases ship different sensors, prefer a release-specific split; recordings of a mixed split publish their missing sensors as empty messages.
Run the ROS node to convert and store the data to rosbags while visualizing it in Rviz.

The Zenseact Open Dataset (ZOD) is a multimodal driving dataset recorded by Zenseact over two years in 14 European countries. Its sensor suite is a single forward-looking 8 MP fisheye camera, three roof-mounted Velodyne lidars (one VLS128 and two VLP16) merged into one point cloud per scan, and an OxTS RT3000 GNSS/IMU. The dataset is released under a permissive license (CC BY-SA 4.0), which allows both research and commercial use.
ZOD is published as three sub-datasets, which are selected together with the version and the split through dataset_split in the form <subset>_<version>_<split>:
| Sub-dataset | Content | Annotations |
|---|---|---|
frames | 100.000 independent keyframes from all over Europe, each with one camera image, one second of surrounding lidar scans in either direction and GNSS/IMU data | Fully annotated |
sequences | 1.473 clips of 20 seconds with the complete sensor suite at 10 Hz | Keyframe (middle frame) only |
drives | 29 clips of a few minutes with the complete sensor suite at 10 Hz | Not annotated |
| Split | Scenes | Samples |
|---|---|---|
frames_full_<train\|val\|all> | 100.000 frames | 1 per frame |
sequences_full_<train\|val\|all> | 1.473 sequences of 20 seconds | approx. 200 per sequence |
drives_full_<train\|val\|all> | 29 drives of a few minutes | approx. 10 per second |
frames_mini_all | 12 frames (10 train, 2 validation) | 12 |
sequences_mini_all | 2 sequences (1 train, 1 validation) | approx. 400 |
drives_mini_all | 2 drives (1 train, 1 validation) | approx. 4.700 |
ZOD calibrates its sensors against an ISO-8855 reference frame at the center of the rear axle at ground level, which is published as base_link.
| Source | Topic | Type | Description |
|---|---|---|---|
| Sensor: Roof Lidars (1x Velodyne VLS128, 2x Velodyne VLP16) | /lidar_01/point_cloud | sensor_msgs/msg/PointCloud2 | Point cloud in the lidar frame (approx. 254.000 points per scan) with float32 fields (x, y, z, intensity), a float64 absolute-seconds timestamp preserving the native per-point timing, and the uint8 diode_index identifying the emitter, and therefore the lidar, a point was measured by. ZOD merges the returns of all three lidars into a single scan. |
| Sensor: Front Camera (8 MP fisheye, 120° HFOV) | /camera_01/image_raw/camera_01/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Anonymized RGB images (height=2168px, width=3848px) with the native Kannala-Brandt calibration, published as the equidistant distortion model. |
| EgoData | /ego_data | perception_msgs/msg/EgoData | Ego-vehicle pose in a local ENU map frame, plus the velocity, acceleration and yaw rate of the high-precision GNSS/IMU interpolated onto the sample's timestamp. |
| Annotation: 3D Lidar Objects | /object_list/lidar_01 | perception_msgs/msg/ObjectList | Annotated 3D objects (HEXAMOTION model) in the lidar_01 frame they are annotated in. |
| Annotation: 3D Camera Objects | /object_list/camera_01 | perception_msgs/msg/ObjectList | The same objects, transformed into the camera_01 frame. |
| Meta Information: Object Annotations | /object_list/lidar_01/meta_info/object_list/camera_01/meta_info | autonomy_datasets_msgs/msg/ObjectListMetaInfo | Annotations without a representation in perception_msgs/msg/Object: original_class, original_subclass, annotation_uuid, unclear, and object_type, occlusion_level, with_rider, emergency, artificial and traffic_content_visible wherever ZOD annotates them. Associated with the object list via the header stamp and the object id. |
| Transformations | /tf, /tf_static | tf2_msgs/msg/TFMessage | Static transformations from the ISO-8855 vehicle frame (base_link) to the sensor frames, and the dynamic pose of base_link in the map frame. |
dataset_split to a frames split to obtain annotated samples only. /clock. Sequences and drives are continuous recordings and are split into scenes of zod_rosbag_duration_seconds instead. perception_msgs/msg/Object and are left out of the object lists. UNKNOWN: ZOD annotates poles, traffic signs, traffic signals, traffic guides and dynamic barriers alongside vehicles and vulnerable road users. perception_msgs/msg/ObjectClassification has no class for them, so they are classified as UNKNOWN, i.e. "definitely none of the other defined classes"; their ZOD class is preserved in original_class and original_subclass. Objects that ZOD flags as unclear are published as UNCLASSIFIED. perception_msgs and are not converted. Radar, which later ZOD releases add for sequences and drives, is not converted either. EgoData reports the dimensions of a large passenger estate car, consistent with the released calibration and with the ego-return box of the development kit.The camera runs at 10.1 Hz and the lidar at 9 Hz, and ZOD ships no synchronization table, so each sample is built from the frames closest in time to the reference sensor, which is the camera because ZOD defines the camera images as its keyframes. Frames without a match within zod_sync_tolerance_seconds are skipped, which typically drops the first sample of a sequence. Point clouds are motion-compensated onto the sample's timestamp, so that lidar, camera and annotations describe the same instant.
The poses ZOD publishes are relative to the first GNSS/IMU sample of a scene, with the x axis along the ego vehicle's heading at that sample. They are rotated by that heading, which is read from the scene's oxts.hdf5, so that map is an ENU frame anchored at that first sample. Scenes of the frames sub-dataset are independent recordings from different places, so their map frames are unrelated to each other.
The dataset requires registration: apply for access to receive a personal download link. Set it via the zod_download_url parameter in params_zenseact_open_dataset.yml or via the ZOD_DOWNLOAD_URL environment variable, and the node downloads and extracts the selected sub-dataset on the first run. Alternatively, download it manually with the CLI of the development kit:
Both ways produce the following folder structure, in which all three sub-datasets live next to each other:
Sub-datasets downloaded into separate directories are picked up as well: an index that is not found in the dataset directory itself is also looked up one level below it, e.g.
$DATASET_DIR/zenseact_open_dataset/frames_mini/trainval-frames-mini.json.
Select the sub-dataset, version and split using dataset_split in params_zenseact_open_dataset.yml, e.g. frames_mini_val, sequences_full_train or drives_mini_all. ZOD provides two anonymizations of its camera images, deep fake anonymization (dnat) and blurring (blur); the frames sub-dataset ships both and is selected via zod_anonymization, while sequences and drives ship the blurred images only.
Run the ROS node to download, convert, and store the data to rosbags while visualizing it in Rviz.

FZI-AURA is a multimodal driving dataset recorded across southern Germany with CoCar NextGen, the research vehicle of the FZI Research Center for Information Technology. It carries the largest lidar suite of any public automated driving dataset: six rotating Ouster lidars (4x OS1-64, 2x OS2-128) and six Aeva Aeries II FMCW lidars give 360° coverage twice over, complemented by eight global-shutter surround-view cameras, up to three Continental ARS 548 RDI radars and an INS. Besides 3D boxes it ships more than 30 billion human-annotated semantic lidar points, more than any other public non-synthetic driving dataset. It is released under a permissive license (CC BY-SA 4.0), which allows both research and commercial use.
| Split | Scenes | Samples |
|---|---|---|
train | 1.979 | approx. 200 per scene |
val | 247 | approx. 200 per scene |
test | 247 | approx. 200 per scene |
all | 2.473 | 493.754 |
Each scene is a self-contained recording of roughly 20 seconds sampled at 10 Hz and becomes one rosbag scene. FZI-AURA annotates 2 Hz keyframes only and releases the sensor payloads of keyframes and non-keyframes as separate download layers, which fzi_aura_samples selects between: keyframes publishes the annotated 2 Hz samples covered by the default download, all publishes the full 10 Hz sample stream.
The sensor suite differs between scenes (4 to 12 lidars, 0 to 3 radars), so a sensor is only mapped to a topic when at least one selected scene holds it. The canonical topic of a sensor is fixed by its position in the suite, so the same sensor always reaches the same topic:
| Topic | Camera | Topic | Lidar | Topic | Radar |
|---|---|---|---|---|---|
camera_01 | front_medium | lidar_01 | top_left | radar_01 | front_left |
camera_02 | front_wide | lidar_02 | top_right | radar_02 | front_right |
camera_03 | front_tele | lidar_03 | front_left | radar_03 | rear_center |
camera_04 | right_forward | lidar_04 | front_right | ||
camera_05 | right_rearward | lidar_05 | rear_left | ||
camera_06 | rear_wide | lidar_06 | rear_right | ||
camera_07 | left_rearward | lidar_07 | aeva_front_center | ||
camera_08 | left_forward | lidar_08 | aeva_front_left | ||
lidar_09 | aeva_front_right | ||||
lidar_10 | aeva_side_left | ||||
lidar_11 | aeva_side_right | ||||
lidar_12 | aeva_rear_center |
FZI-AURA calibrates its sensors against base_link, which already follows the ROS convention (x forward, y left, z up) and is published unchanged. Sensor frames are published as <modality>_<sensor id>, e.g. lidar_top_left or radar_front_left, because a bare sensor ID is not unique across modalities.
| Source | Topic | Type | Description |
|---|---|---|---|
| Sensor: Ouster Lidars (4x OS1-64, 2x OS2-128) | /lidar_01/point_cloud.../lidar_06/point_cloud | sensor_msgs/msg/PointCloud2 | Point cloud in the lidar frame (approx. 60.000 points per scan) with the native fields (x, y, z, t, reflectivity, ring, range). Motion-compensated onto the end of the sweep by default, selectable via fzi_aura_lidar_stage. |
| Sensor: Aeva Aeries II FMCW Lidars | /lidar_07/point_cloud.../lidar_12/point_cloud | sensor_msgs/msg/PointCloud2 | Point cloud in the lidar frame with the native fields of the FMCW sensor, including the per-point Doppler velocity measured along the beam and an intensity next to the reflectivity. The raw and the motion-compensated stage carry different fields. |
| Sensor: Surround-View Cameras (8x global shutter) | /camera_01/image_raw.../camera_08/image_raw/camera_0X/camera_info | sensor_msgs/msg/Imagesensor_msgs/msg/CameraInfo | Anonymized RGB images, 1920x1200px for the six narrow cameras and 2592x2048px for front_wide and rear_wide, scalable via fzi_aura_image_scale. The images are rectified and are published with their projection matrix and zero distortion coefficients. |
| Sensor: Radars (up to 3x Continental ARS 548 RDI) | /radar_01/point_cloud.../radar_03/point_cloud | sensor_msgs/msg/PointCloud2 | Radar detections in the radar frame with the fields (x, y, z, rcs, elevation_angle) and, where the sensor reports it, the radial velocity range_rate. |
| Annotation: Semantic Lidar Labels | /lidar_01/point_cloud.../lidar_06/point_cloud | sensor_msgs/msg/PointCloud2 | Per-point semantic_id and instance_id fields of the cloud they annotate, added wherever the scene is semantically labeled. Controlled via fzi_aura_publish_semantic_labels. |
| EgoData | /ego_data | perception_msgs/msg/EgoData | Ego-vehicle dimensions and dynamics state (EGO model) in an ENU map frame. Pose from the dataset's ego poses, velocity, acceleration, yaw rate, steering angle, standstill flag and turn indicator from the INS and the vehicle's CAN bus. |
| Annotation: 3D Lidar Objects | /object_list/lidar_01 | perception_msgs/msg/ObjectList | Annotated 3D objects (HEXAMOTION model) in the frame of the reference lidar, restricted to the objects holding at least one point of that lidar. |
| Annotation: 3D Vehicle-Frame Objects | /object_list/base_link | perception_msgs/msg/ObjectList | The canonical annotation: every 3D object of the keyframe in the base_link frame, including objects no single lidar sees points of. |
| Meta Information: Object Annotations | /object_list/lidar_01/meta_info/object_list/base_link/meta_info | autonomy_datasets_msgs/msg/ObjectListMetaInfo | Annotations without a representation in perception_msgs/msg/Object: original_class, the object_id the dataset tracks an object under within a scene, and the sensor_id of the per-sensor object list. Associated with the object list via the header stamp and the object id. |
| Transformations | /tf, /tf_static | tf2_msgs/msg/TFMessage | Static transformations from the vehicle frame (base_link) to every calibrated sensor frame, and the dynamic pose of base_link in the map frame. |
| Map | map_contents (parameter) | string | Lanelet2 map (OSM XML) of the roads around the current scene, generated from OpenStreetMap. Updated on every scene change, analogous to lanelet2_map_server. |
FZI-AURA ships no map, so the Lanelet2 map of a scene is generated from the OpenStreetMap roads within 200 m of its GNSS track, which are fetched from the Overpass API configured via fzi_aura_overpass_url. The map generation requires internet access and is controlled via the publish_lanelet2_map (enable/disable) and fzi_aura_lanelet2_lane_width (assumed lane width in meters) parameters. A scene that cannot be fetched is published without a map. The map is stored next to the rosbag data of its scene like the nuScenes map, so replaying a rosbag does not fetch it again.
Each road becomes one road lanelet per lane (highway for motorways, play_street for living streets), laid out around its OSM centerline from the lanes, lanes:forward, lanes:backward and oneway tags for right-hand traffic; the lane boundaries are synthesized by offsetting the centerline by multiples of the lane width. Crossings mapped as footway=crossing become crosswalk lanelets. The map is georeferenced into the map frame by the GNSS fixes of the scene's samples, which agree with the ego poses to a few decimeters.
fzi_aura_samples: keyframes publishes the annotated samples only. fzi_aura_samples: all, the samples between two keyframes are published without sensor data unless the camera_nonkeyframes, lidar_raw_nonkeyframes and radar_nonkeyframes layers were downloaded as well. Motion-compensated lidar is not released for non-keyframes at all. map frame is aligned with the INS attitude: The frame the released ego poses are expressed in carries an arbitrary orientation per recording that is neither gravity-aligned nor north-referenced — the poses of a scene can hold a roll of more than ten degrees while the vehicle drives level. The INS state in the vehicle signals is a proper east-north-up attitude, so the poses are rotated by the offset between the two at the first sample of a scene, which yields an ENU-aligned map. The residual drift of the released odometry stays below about two degrees over a scene. Scenes without vehicle signals keep the native orientation of the released poses. bicycle, motorcycle, portable) and as the vehicle together with its rider (bicyclist, motorcyclist, portable-rider). perception_msgs/msg/ObjectClassification defines BICYCLE, MOTORCYCLE and MICRO as covering the vehicle and its rider, so both spellings map to the same class; a rider box holds the person alone and is published as VRU. Objects of the dynamic class, which collects movable objects that fit none of the other classes, are published as UNKNOWN. The dataset's own class is preserved in original_class.The dataset is available on Hugging Face: after log in, the node will download the selected scenes on the first run via the FZI-AURA SDK downloader.
Alternatively, download the data manually with the CLI of the SDK:
Both ways produce the following folder structure:
Select the split and the scenes using dataset_split and fzi_aura_scenes in params_fzi_aura.yml.
Run the ROS node to download, convert, and store the data to rosbags while visualizing it in Rviz.
Custom datasets according to your needs and suitable for commercial use are available via an expanding network of partners on request, for example:
perception_msgs/msg/Object (e.g. the dataset's original class name) as autonomy_datasets_msgs/msg/ObjectListMetaInfo on the object list's meta_info topic, using the helpers in meta_info.py.