Andreas Geiger

Publications of Gege Gao

Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories
M. Kocabas, G. Gao, B. Schölkopf and A. Geiger
European Conference on Computer Vision (ECCV), 2026
Abstract: Diffusion and flow-based generative models produce strong images, yet their controllability remains largely endpoint-centric: users specify conditions and receive final outputs, while the intermediate generative dynamics remain hidden. Recent methods have begun to exploit generation order and process decomposition to improve sample quality, but still treat intermediate states as internal computation rather than objects for interaction. We propose Trajectory Forcing (TF), a trajectory-centric framework that makes the generation path explicit, semantic, and editable. TF organizes synthesis as a sequence of semantically structured stages, progressing from global layout to object-, part-, and detail-level representations. Each stage produces a decodable latent state that can be inspected, evaluated, and locally edited before the next stage begins. To instantiate this path, we derive coarse-to-fine teacher hierarchies by clustering pretrained visual representations such as DINOv2, and train a hierarchy-conditioned one-step flow-matching model at each level. We further introduce trajectory-aware metrics that measure structural consistency and local controllability beyond endpoint quality metrics such as FID. Experiments show that TF achieves competitive sample quality while exposing coherent intermediate states and supporting localized edits across semantic levels. By shifting the focus from final images to the generative path itself, TF opens a route toward controllable, trajectory-aware image synthesis.
Latex Bibtex Citation:
@inproceedings{Kocabas2026ECCV,
  author = {Merve Kocabas and Gege Gao and Bernhard Schölkopf and Andreas Geiger},
  title = {Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year = {2026}
}
Echoes of the Prior: A Computational Phenomenology of Forgetting
G. Gao, B. Schölkopf and A. Geiger
International Conference on Computer Graphics and Interactive Techniques (SIGGRAPH), 2026
Abstract: Memory is not merely the storage of data; it is the scaffolding of reality. When biological memory fades, the world does not simply turn black; it regresses into an unrecognizable chaos. Echoes of the Prior is an interactive installation that attempts to visualize this subjective phenomenology of forgetting. By inducing controlled synaptic decay within a Feed-Forward 3D Reconstruction model, we create an artistic analogy for the erosion of the brain's predictive priors. We position the Neural Network not as a tool for engineering, but as a cognitive proxy - a silicon brain whose structural degeneration evokes the disorienting, poetic, and terrifying experience of losing one's grip on the world. Ultimately, we offer this framework as a catalyst, inviting the wider community to explore the uncharted potential of neuromorphic aesthetics in visualizing the fragility of intelligence. Interactive demo see https://decart-4d.github.io/.
Latex Bibtex Citation:
@inproceedings{Gao2026SIGGRAPH,
  author = {Gege Gao and Bernhard Schölkopf and Andreas Geiger},
  title = {Echoes of the Prior: A Computational Phenomenology of Forgetting},
  booktitle = {International Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)},
  year = {2026}
}
GraphDreamer: Compositional 3D Scene Synthesis from Scene Graphs
G. Gao, W. Liu, A. Chen, A. Geiger and B. Schölkopf
Conference on Computer Vision and Pattern Recognition (CVPR), 2024
Abstract: As pretrained text-to-image diffusion models become increasingly powerful, recent efforts have been made to distill knowledge from these text-to-image pretrained models for optimizing a text-guided 3D model. Most of the existing methods generate a holistic 3D model from a plain text input. This can be problematic when the text describes a complex scene with multiple objects, because the vectorized text embeddings are inherently unable to capture a complex description with multiple entities and relationships. Holistic 3D modeling of the entire scene further prevents accurate grounding of text entities and concepts. To address this limitation, we propose GraphDreamer, a novel framework to generate compositional 3D scenes from scene graphs, where objects are represented as nodes and their interactions as edges. By exploiting node and edge information in scene graphs, our method makes better use of the pretrained text-to-image diffusion model and is able to fully disentangle different objects without image-level supervision. To facilitate modeling of object-wise relationships, we use signed distance fields as representation and impose a constraint to avoid inter-penetration of objects. To avoid manual scene graph creation, we design a text prompt for ChatGPT to generate scene graphs based on text inputs. We conduct both qualitative and quantitative experiments to validate the effectiveness of GraphDreamer in generating high-fidelity compositional 3D scenes with disentangled object entities.
Latex Bibtex Citation:
@inproceedings{Gao2024CVPR,
  author = {Gege Gao and Weiyang Liu and Anpei Chen and Andreas Geiger and Bernhard Schölkopf},
  title = {GraphDreamer: Compositional 3D Scene Synthesis from Scene Graphs},
  booktitle = {Conference on Computer Vision and Pattern Recognition (CVPR)},
  year = {2024}
}


eXTReMe Tracker