Code and data for "Situation Graph Prediction for User Perspective Modeling".
Situation Graph Prediction (SGP) frames perspective modeling as an inverse inference problem: reconstructing structured, ontology-aligned representations of latent perspective from observable multimodal artifacts.
sgp/
├── README.md
├── LICENSE
├── requirements.txt # Python dependencies
├── data/
│ ├── situations/ # 75 instances (event_1..75); each = one (X_t, G_t) pair
│ │ └── event_N/ # x1_*.mp3 (audio), x2_*.png (image), x3_*.txt (text), label_*.txt (G_t)
│ ├── schema/ # ontology: node kinds, edge types, A(p) constraints (ontology.py)
│ ├── dataset_metadata.csv # per-instance metadata (artifact paths + triples)
│ └── dataset_metadata_vocabulary.csv # indexed vocabulary of predicates/entities
├── evaluation/
│ ├── run_experiment.py # main runner: zero-shot and RA-ICL, 5-fold CV
│ ├── decompose.py # modality decomposition (audio transcription via Whisper)
│ ├── retrieval.py # RA-ICL retrieval (text-embedding-3-large, k=3)
│ ├── prompt_builder.py # zero-shot / RA-ICL prompt construction
│ ├── metrics.py # PVR, strict/soft F1, latent-vs-surface, entropy-normalized F1*
│ ├── dataset_manager.py # dataset loading + stratified k-fold splits
│ └── models/ # per-model handlers: gpt4o, gemini, claude (+ base)
├── results/
│ ├── raw/{gpt4o,gemini,claude}-run{1,2,3}/ # per-run, per-fold, per-model raw predictions
│ ├── aggregated.csv # the numbers behind Table 1 (generated by make_table1.py)
│ └── make_table1.py # rebuilds Table 1 from results/raw/ (mean ± within-run across-fold std)
└── supplementary/ # taxonomies, per-model P/R, entropy derivation, qualitative examples
Requires Python 3.10+.
pip install -r requirements.txt
# Set the API keys for whichever model(s) you intend to run
export OPENAI_API_KEY="your-key" # GPT-4o + text-embedding-3-large + Whisper (Claude audio)
export GOOGLE_API_KEY="your-key" # Gemini 2.5 Flash
export ANTHROPIC_API_KEY="your-key" # Claude Sonnet 4The pilot dataset (data/situations/) contains 75 situation instances and 225 synthetic artifacts (audio, image, text per instance). All instances center on a single fictional persona, Elise Navarro, a 28-year-old professional in Toronto, across a 2021–2025 (~60-month) timeline spanning four domains: professional developments, personal/lifestyle changes, health and physical milestones, and social/relational experiences. Modalities are primarily text, with image and audio subsets. Each instance pairs a set of artifacts X_t with a ground-truth situation graph G_t (label_*.txt).
All data is synthetic. No real user data was collected or used. The persona and all artifacts were produced by the structure-first generation pipeline and curated for consistency.
The pipeline maps raw artifacts to predicted triples in three stages — modality decomposition, then per-mode inference (zero-shot and RA-ICL), then aggregation. Run from the repository root:
# 1. Decomposition (Phase A) + zero-shot inference.
# The first run decomposes every artifact — image -> VLM description,
# audio -> transcript, text -> passthrough — and caches the result to
# <output_dir>/description_cache.json. There is no separate decompose command.
python evaluation/run_experiment.py --model gpt4o --cross_validate --mode zero_shot \
--output_dir results/raw/gpt4o-run1
# 2. RA-ICL inference (retrieval with text-embedding-3-large, k=3).
# Reuses the cached decompositions from step 1.
python evaluation/run_experiment.py --model gpt4o --cross_validate --mode few_shot \
--output_dir results/raw/gpt4o-run1
# (Steps 1–2 can be combined with `--compare`, and repeated for
# --model gemini and --model claude, and for run1/run2/run3.)
# 3. Rebuild Table 1 from the raw predictions.
python results/make_table1.py # writes results/aggregated.csvCaveat: the study evaluates closed-weight API models (gpt-4o, gemini-2.5-flash, claude-sonnet-4-20250514); these versions may be updated or deprecated, so fresh runs may not reproduce the paper's numbers exactly. The original per-run, per-fold predictions are preserved under results/raw/ as the durable evidence behind Table 1 — make_table1.py recomputes the table from them without re-calling any API.
If you use this code or dataset, please cite our paper:
@misc{shin2026sgp,
author = {Shin, Jisung and Platnick, Daniel and Alirezaie, Marjan and Rahnama, Hossein},
title = {Situation Graph Prediction for User Perspective Modeling},
year = {2026},
eprint = {2602.13319},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
doi = {10.48550/arXiv.2602.13319},
url = {https://arxiv.org/abs/2602.13319}
}MIT License — see LICENSE.