💡 This post is initially focused on interpretability for multimodal models, while later a lot of papers in other fields are included, just for convenience.

Methods

Interpretability for MLLMs

Interpretability for Diffusion Models

Other fields of MLLMs

Datasets & Benchmarks

general

small-object

Coding & GUI

spatial

video

hallucination

Training

Learning Dynamics

Optimization

Data

Infra

Models

LLM

self-supervised learning

MLLM

generative models

world models