HuyMichel's picture
|
download
raw
3.44 kB

EventActivityNet Dataset Format

Representations

EventActivityNet stores one HDF5 file per video for each supported bin count.

Bin count B Public payload Event shape
5 data_5bin/ (T5, 5, H, W)
9 data_9bin/ (T9, 9, H, W)

File size varies substantially with duration and spatial resolution.

HDF5 Schema

Every HDF5 file has exactly these root datasets:

events
voxel_event_start
voxel_event_count

events

Property Value
Shape (T_B, B, H, W)
Dtype int16
Compression gzip, level 4
Shuffle enabled
Chunking (1, B, min(H, 256), min(W, 256))

voxel_event_start

Property Value
Shape (T_B,)
Dtype int64
Compression LZF
Shuffle enabled
Chunking (1024,)

voxel_event_count

Property Value
Shape (T_B,)
Dtype int32
Compression LZF
Shuffle enabled
Chunking (1024,)

Required root attributes are fps, height, width, num_bins, and interpolate_bins. num_bins is 5 or 9 and matches events.shape[1].

Transition Grouping

Construction follows decoded frame order. For N source frames there are N - 1 adjacent-frame transitions. Transition index e corresponds to source frames (e, e + 1).

For bin count B:

T_B = ceil((N - 1) / B)
voxel_event_start[t] = B * t
voxel_event_count[t] = min(B, N - 1 - B * t)

events[t] groups transitions in the half-open range [B*t, min(B*t + B, N - 1)). The associated source-frame interval is [B*t, min(B*t + B, N - 1)]. The final group can contain fewer than B valid transitions; unused bins are zero-filled.

Timing

The HDF5 fps attribute is source-frame FPS stored as a float. Use the released fps_num and fps_den fields for reproducible conversion. Approximate voxel times are:

start_seconds = B * t * fps_den / fps_num
end_seconds = min(B * t + B, N - 1) * fps_den / fps_num

Per-frame presentation timestamps are not consumed. These conversions are therefore approximate for within-video variable-frame-rate streams. Do not assume fixed 25 fps or 240 fps, and do not use t / fps as voxel time.

For a caption/action interval [start_seconds, end_seconds]:

start_frame = floor(start_seconds * fps_num / fps_den)
end_frame = ceil(end_seconds * fps_num / fps_den)
t_start = max(0, floor(start_frame / B))
t_end_exclusive = min(T_B, ceil(end_frame / B))

Use [t_start, t_end_exclusive) for Python slicing.

Memory-Safe Loading

import h5py

with h5py.File("v_example.h5", "r") as f:
    events = f["events"]
    starts = f["voxel_event_start"]
    counts = f["voxel_event_count"]

    B = int(f.attrs["num_bins"])
    print(events.shape)  # (T_B, B, H, W)
    print(events.dtype)  # int16
    print(starts.dtype)  # int64
    print(counts.dtype)  # int32

    selected = events[10:18]  # reads only the selected temporal range

Avoid loading complete event tensors unless sufficient memory is available.

Shards and Metadata

Each representation has 157 train shards, 62 validation shards, and 3,263 HDF5 members. Representation-specific manifests and checksums are under:

metadata/5bin/
metadata/9bin/

Shared source/timing metadata is in metadata/video_metadata.jsonl, with representation-specific tensor fields nested under representations.

Xet Storage Details

Size:
3.44 kB
·
Xet hash:
12cf90a158b988a14017ddcafe6ddc827666aab928de6ea94d61d3b516a9b7db

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.