Vision-Language Models for Image-Based Dietary Assessment: A Benchmark of Accuracy, Cost, and Prompt Strategies Across Ten Models

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Dietary assessment is the cornerstone of clinical management and research studies evaluating diet and health. Traditional methods such as food diaries and 24-hour recalls can be burdensome, prone to recall bias, and difficult to adhere to. Image-based dietary assessment using vision-language models (VLMs) offers a potential solution.

Objective

Our goal was to benchmark state-of-the-art VLMs for automated food recognition, weight estimation, and calorie estimation using Google’s Nutrition5k dataset.

Methods

We evaluated 3,229 food images using ten approaches: proprietary VLMs (Gemini 2.0 Flash, 2.5 Flash, 3.0 Flash, and 3.1 Flash Lite; GPT- 4o, GPT-4o-mini, and GPT-5 Mini; and Claude Haiku 4.5), an open-source VLM (Qwen2-VL-7B), and a commercial food recognition API (FatSecret). We assessed calorie and weight estimation using Lin’s Concordance Correlation Coefficient (CCC) and component detection using Jaccard similarity.

Results

Gemini 3.0 Flash achieved the best calorie estimation (CCC 0.767, MAE 80.7 kcal), while Gemini 3.1 Flash Lite offered very comparable accuracy (CCC 0.754) with the highest ingredient recognition (Jaccard 0.655) at the lowest cost among top-performing models ($0.59/1K images). Among earlier-generation models, Gemini 2.0 Flash remained competitive (CCC 0.742, Jaccard 0.621) at a fraction of the cost ($0.10/1K images). A human validation study in which four annotators reviewed 440 images revealed systematic omissions in the original Nutrition5k labels. After correction, the extrapolated ingredient-overlap score for Gemini 2.0 Flash increased from 0.62 to an estimated 0.82, suggesting that raw Jaccard scores substantially underestimate true model performance.

Conclusions

Current VLMs can perform automated dietary assessment with reasonable accuracy from single overhead photographs. Our results inform model selection for dietary assessment applications and highlight remaining challenges in calorie estimation and component detection for complex, multi-item meals.

Article activity feed