Meta has introduced a new benchmark called GAMUT to measure how factually complete AI models are. The benchmark tests models on 1,813 multimodal questions. The best-performing model, Gemini 3.1 Pro, scored just 58.7%.
What GAMUT measures
GAMUT stands for something Meta hasn't fully spelled out, but the goal is clear: evaluate whether AI systems give complete factual answers, not just plausible ones. Most benchmarks test accuracy on narrow tasks. GAMUT instead checks if a model covers all relevant facts when responding to a question. That means it looks for omissions, not just errors.
The questions are multimodal, meaning they combine text and images. For example, a question might include a photo of a historical event and ask for details about it. The model must pull facts from both the image and its training data to give a full answer.
The 1,813 questions
Meta built the benchmark around 1,813 questions. Each question has a set of required facts that a complete answer should include. The model's response is scored on how many of those facts it covers. This is different from a simple right-or-wrong test. A model could get every fact right but still miss half of them — and that would be a low GAMUT score.
The questions span topics like history, science, geography, and current events. They require models to pull from both text and image inputs, which adds complexity.
Why the top score matters
Gemini 3.1 Pro's 58.4% is the highest among tested models, but it still means the model misses more than 40% of relevant facts. That's a big gap. For applications where completeness matters — like medical advice, legal research, or news summarization — missing facts can be dangerous.
Meta hasn't released scores for its own models, like Llama, on GAMUT. The company also hasn't said whether it will make the benchmark public for other researchers to use. That leaves open questions about how widely GAMUT will be adopted.
GAMUT is one of several recent efforts to push AI evaluation beyond simple accuracy. Google, OpenAI, and others have their own benchmarks for truthfulness and hallucination. But GAMUT focuses specifically on completeness, which is a different angle.
Meta has not announced plans to update GAMUT or release further results. For now, the benchmark serves as a reminder that even the best AI models have significant gaps in what they know — and what they tell us.




