arXiv Open Access 2025

Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs

Junxian Li Xinyue Xu Sai Ma Di Zhang Sichao Li

Lihat Sumber

Abstrak

Multimodal Large Language Models (MLLMs) frequently suffer from unfaithfulness, generating reasoning chains that drift from visual evidence or contradict final predictions. We propose Faithful-First Reasoning, Planning, and Acting (RPA) framework in which FaithEvi provides step-wise and chain-level supervision by evaluating the faithfulness of intermediate reasoning, and FaithAct uses these signals to plan and execute faithfulness-aware actions during inference. Experiments across multiple multimodal reasoning benchmarks show that faithful-first RPA improves perceptual faithfulness by up to 24% over prompt-based and tool-augmented reasoning frameworks, without degrading task accuracy. Our analysis shows that treating faithfulness as a guiding principle perceptually faithful reasoning trajectories and mitigates hallucination behavior. This work thereby establishes a unified framework for both evaluating and enforcing faithfulness in multimodal reasoning. Code will be released upon acceptance.

Topik & Kata Kunci

cs.AI

Penulis (5)

Junxian Li

Xinyue Xu

Sai Ma

Di Zhang

Sichao Li

Format Sitasi

APA MLA BibTeX

Li, J., Xu, X., Ma, S., Zhang, D., Li, S. (2025). Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs. https://arxiv.org/abs/2511.08409

Akses Cepat

Lihat di Sumber

Informasi Jurnal

Tahun Terbit: 2025
Bahasa: en
Sumber Database: arXiv
Akses: Open Access ✓