EgoRecall

EgoRecall

3D Visual Grounding from Streaming Egocentric Observations

Hou In Ivan Tam,  Manolis Savva

Simon Fraser University

Unlike conventional 3DVG, which assumes a complete scene is available in advance, our task grounds language using only the partial experience accumulated by an agent over time. Left: An agent perceives a scene through a stream of egocentric views. Right: At query time τ it receives a referring expression that may depend on when objects were observed and from where, and it must find the referred objects in 3D, even when they are no longer visible.

Abstract

We introduce streaming egocentric 3D visual grounding, where an agent resolves referring expressions using only the partial observation history available at query time. Unlike conventional 3DVG, which assumes a complete scene representation, queries may depend on when and how objects were encountered as well as their spatial relationships.

We present EgoRecall, the first benchmark for this task, comprising 2.25M timestamped queries generated by a compositional DSL over observation history, egocentric relationships, and allocentric structure. Queries are defined with respect to the observations available at their timestamps and may refer to either currently visible or previously observed objects, requiring persistent scene memory.

We evaluate VLM-, 3D-, and memory-based methods under protocols that separate perception from grounding. The task is challenging with our strongest system achieving only 14.1 F1. Our experiments show that modeling observation history and camera-relative relationships improves performance, but targets no longer visible remain particularly challenging.

Grounding from what you have already seen

Consider a user wearing smart glasses who walks through an unfamiliar office and then asks for “the cabinet to the right of the first sink I encountered”. Resolving this query requires finding the first sink, recovering the viewpoint from which it was seen, and then locating the cabinet to its right, although both objects may have left the view long ago. References in this setting draw on three axes.

A trajectory through a room with timers where objects were seen.

Observation history

When and how objects were encountered: their order and recency.

Example query in the dataset

the earliest heater

first_seen(heater)

Asked at 3:02
Target out of view for 109 s

A person looking at two boxes: one is left of the other from their viewpoint.

Egocentric relationships

Relations to the observer: what is in view, how far, and which side.

Example query in the dataset

the cabinet closest to me

nearest_to_cam(cabinet)

Asked at 2:32
Target out of view for 52 s

Two monitors next to each other on a desk.

Allocentric structure

Relations among objects that do not depend on the viewpoint.

Example query in the dataset

the paper roll nearest to the first shoes that appeared

nearest(first_seen(shoes), paper roll)

Asked at 3:00
Target out of view for 46 s

The example queries are from test scene 61adeff7d5 (queries 10551, 9128, and 10456); click one to open it in the scene explorer below.

The EgoRecall dataset

We introduce EgoRecall, the first dataset and benchmark for this task. Our queries are generated by executable programs over what the camera has seen, so every answer is exact at the moment it is asked. Most queries refer to objects that are no longer in view.

2.25Mqueries
183ScanNet++ scenes
12.1 hof video
21,331objects
1,366object categories
22program operators

Record visibility

Each ScanNet++ capture is replayed at 6 FPS. At every camera pose, the annotated mesh is rendered and checked against the sensor depth to record which objects are visible.

Run programs

A query is a program over 22 operators, nested up to two levels. Running it at a frame uses only the objects seen by then, which gives its exact targets.

Emit and balance

Along each video, programs run every 2 s and are turned into text with templates. The 9.9M generated queries are then balanced across scenes, operator families, and target visibility into the final set of 2.25M queries.

left_of(first_seen(box), floor mat)

“the floor mat to the left of the earliest boxes” Asked at 1:30 in test scene 61adeff7d5 · target out of view for 15 s

Are the targets visible when the query is asked?

Some or all visibleNone visible, by how long ago a target was last seen →

In 80% of queries, at least one target is not visible when the query is asked and has to be remembered. Among queries with no visible target, 46% refer to one last seen more than 10 s earlier, and 15% refer to one last seen more than a minute earlier. The dataset is balanced to these six proportions by design.

Which reference axes a query uses

82% of queries compose two or more operators, and 48% combine two of observation history, egocentric relationships, and allocentric structure in the same expression. Queries have 1.6 targets on average, and 42% refer to more than one object.

The 30 most frequent target categories, of 1,366

Splits and evaluation stages

SplitScenesVideoObjectsQueriesStages
Train1006.4 h11,5821,202,906–
Validation130.8 h1,616165,34083
Test704.8 h8,133879,994440

Scenes do not overlap across splits. Validation and test queries are divided into stages of 2,000, drawn so that stages 1 to k always form a proportional random sample; the paper evaluates test stages 1–5 (10,000 queries). See the dataset card for more details.

Explore a scene

Replay a test scene and pick a query to see what it refers to and how each method answered.

0:00
  • Camera now
  • Path so far
  • In view
  • Remembered
  • Target
  • Anchor
  • Wrong pick
Loading the scene…
0:00 / 4:06

Featured queries

Test scene 61adeff7d5 from ScanNet++. Drag to rotate the view and scroll to zoom.

Results

We adapt five methods from four families and evaluate them on 10,000 test queries. To answer a query, a method can use only what has been observed up to the moment the query is asked.

VLM pointing

Qwen 3.8-27B

Chooses a frame among up to 64 RGB frames and points at each target.

Memory-free 3D grounding

Qwen3D-3B

Rebuilds a point cloud from up to 128 RGB-D frames for every query.

Scene-graph memory

DAAAM, DirectMe

Build a scene graph as frames arrive; a VLM selects its nodes.

Object-centric memory

MTU3D, EgoR-MTU3D

Keeps a memory bank of objects; EgoR-MTU3D adds history, recency, and viewpoint encoders.

Finding 1

The task is far from solved

The best system, EgoR-MTU3D, reaches only 14.1 F1full.

Finding 2

Remembered targets are harder

Methods score lower when all targets must be remembered than when all are visible.

Finding 3

Perception is a bottleneck

Memory-based methods depend heavily on perception. With its own estimates of camera pose and depth, DirectMe keeps a target in its scene graph for only 31% of queries; given ground-truth poses and sensor depth, it keeps one for 73%, and its F1full rises from 0.5 to 8.7.

DirectMe withCov%F1full
Estimated pose and depth31.4%0.5
Ground-truth pose62.6%6.7
Ground-truth pose and sensor depth73.0%8.7
Finding 4

Knowing when and from where helps

EgoR-MTU3D adds three encoders to MTU3D's object memory: history (when an object was first seen and how often), recency (whether it is in view and when it was last seen), and viewpoint (its direction and distance from the camera). Each raises F1sel, and all three together raise it most.

Method Pose / Depth Overall F1sel by operator family
Cov% F1sel F1full Top-1 tempn=5,203 visn=2,031 idn=1,544 egoDn=2,853 egoZn=1,889 latn=1,364 proxn=752 supn=1,868 setn=2,347
VLM pointing
Qwen 3.8-27B†– / –99.112.212.113.211.816.417.411.011.17.913.69.218.8
Memory-free 3D grounding
Qwen3D-3BGT / sensor95.910.09.612.09.011.112.010.211.59.213.08.110.3
Scene-graph memory
DAAAMGT / sensor77.710.47.911.69.611.413.010.811.96.911.99.011.3
DirectMeest. / est.31.41.70.50.71.72.31.31.31.31.62.72.21.5
DirectMe-GT-PGT / est.62.611.56.78.410.014.614.411.311.48.69.79.413.0
DirectMe-GT-PDGT / sensor73.012.78.710.911.115.215.313.511.710.214.08.014.9
Object-centric memory
MTU3D zero-shotGT / sensor68.32.81.810.22.12.94.62.12.11.52.22.12.2
MTU3D fine-tunedGT / sensor68.315.910.117.914.519.222.615.616.912.215.610.620.6
+ history encoderGT / sensor68.317.811.318.116.321.424.517.719.713.015.811.822.4
+ recency encoderGT / sensor68.318.011.518.916.224.223.316.420.512.916.811.521.8
+ viewpoint encoderGT / sensor68.319.712.620.917.324.425.320.624.812.320.211.723.4
EgoR-MTU3DGT / sensor68.322.214.122.520.027.927.023.027.415.120.713.826.5

Grounding performance on 10,000 test queries. Full system F1 (F1full) evaluates the complete pipeline, whereas Selection F1 (F1sel) conditions on target availability; Cov% reports this availability. Operator-family columns report F1sel for temporal (temp), visibility (vis), identity (id), egocentric distance (egoD), egocentric depth (egoZ), lateral (lat), proximity (prox), relational-superlative (sup), and set-operation (set) queries. EgoR-MTU3D performs best, reaching 22.2 F1sel and 14.1 F1full, although its 68.3% coverage highlights incomplete object memory as a major bottleneck. † The VLM uses a semi-oracle readout based on GT instance maps.

Qualitative results: top-down maps with ground-truth and predicted boxes for three queries whose targets are out of view.
Qualitative examples. Green boxes mark ground-truth or correct predictions, red boxes mark incorrect predictions, and blue annotations show ground-truth trajectories or relational anchors; all three example targets are out of view at query time τ. Top: memory-based methods—EgoR-MTU3D, DAAAM, and DirectMe-GT-P—correctly recover the target for a compositional query over a 414-second observation history, while methods without persistent object memory fail. Bottom left: viewpoint encoding enables MTU3D to ground the remembered target table, whereas Qwen and fine-tuned MTU3D without the viewpoint encoder select distractors (other visible tables). Bottom right: DirectMe's estimated trajectory drifts from GT, so it misgrounds the pillow, while DirectMe-GT-P yields the correct target.

Demo on Aria glasses

We run EgoR-MTU3D on a real recording from Project Aria glasses.

The user wears the glasses for 24 minutes while moving through a multi-room apartment, and EgoR-MTU3D maps what they see into over 2,000 object tracks. From another room, 76 s after the kettle left the view, the user asks for “the kettle I used to boil water”. EgoR-MTU3D grounds the kettle in 3D, which can be used to guide the user back to it. The end of the clip shows what this guidance could look like: a route to the kettle on the map, and an arrow in the user's view.

Get started

EgoRecall's annotations are on Hugging Face. The frames, depth, and camera poses come from ScanNet++, which you download yourself and prepare with our code. See the README for setup and the dataset card for the data format.

Download the annotations

Request access on Hugging Face, then download the queries, their answers, and the object visibility records (about 110 MB).

hf auth login
hf download 3dlg-hcvc/EgoRecall --repo-type dataset \
  --local-dir /path/to/EgoRecall_hf

Get ScanNet++

Request access to ScanNet++ v2 and download its iPhone captures and object annotations: about 117 GB for the 70 test scenes, or 292 GB for all 183 scenes.

Prepare and load

Install the code and set the paths in configs/paths.toml. Preparing the scenes makes them ready for our Python API, which loads each query sample along with its observation history.

egorecall-prepare --config configs/paths.toml --split test
paths = DatasetPaths.from_toml(Path("configs/paths.toml"))

# Test stages 1-5: the 10,000 queries evaluated in the paper
dataset = EgoRecallDataset(paths, split="test", stages="1:5")
scene_id, query_idx = dataset.annotations.query_keys[0]
with dataset.open_scene(scene_id) as scene_data:
    sample = scene_data.query(query_idx)

Citation

If you find EgoRecall helpful in your research, please cite our work. EgoRecall is built on ScanNet++, so please cite it as well.

@article{tam2026egorecall,
  title={{EgoRecall}: {3D} Visual Grounding from Streaming Egocentric Observations},
  author={Tam, Hou In Ivan and Savva, Manolis},
  journal={alphaXiv preprint},
  year={2026}
}

Acknowledgements

This work was funded in part by a Canada Research Chair, NSERC Discovery Grants, and enabled by support from the Digital Research Alliance of Canada, and an NVIDIA Academic Grant Award. We thank Austin T. Wang, Denys Iliash, and Weikun Peng for helpful feedback and discussions.