Summary
This arXiv paper (2607.09623) by Nirjhar Das and Md. Al-Mamun Provath describes a submission to the QANTA 2026 shared challenge at the EMM-QA workshop, ICML 2026. The authors developed a task-specific dual-agent architecture for multimodal question answering. The Tossup agent uses a GPT-4o-mini-class model with confidence-calibrated answering and domain-specific numerical reasoning strategies to reduce overconfident predictions on isolated quantitative clues. The Bonus agent uses a GPT-4o-class model with clue-aware reasoning, structured relational reasoning, and multimodal evidence integration. On the leaderboard, the system achieved the highest overall score of 0.402, comprising a Tossup score of 0.238 and a Bonus Effect of 0.164. The results demonstrate that lightweight, task-specific reasoning strategies can deliver strong performance on resource-constrained multimodal QA benchmarks.
Paper Overview
Field: NLP/AI
Authors: Nirjhar Das, Md. Al-Mamun Provath
Published: 2026-07-10
arXiv: 2607.09623
Abstract
This paper presents a submission to the QANTA 2026 shared challenge, hosted at the EMM-QA workshop at ICML 2026. The authors developed a task-specific dual-agent architecture for multimodal question answering:
Tossup Agent
- Uses a GPT-4o-mini-class model
- Employs confidence-calibrated answering
- Applies domain-specific numerical reasoning strategies
- Reduces overconfident predictions on isolated quantitative clues
Bonus Agent
- Uses a GPT-4o-class model
- Features clue-aware reasoning
- Performs structured relational reasoning
- Integrates multimodal evidence
Results
The system achieved the highest overall score on the leaderboard at 0.402:
| Metric | Score |
|---|---|
| Total | 0.402 |
| Tossup | 0.238 |
| Bonus Effect | 0.164 |
Conclusion
The results demonstrate that lightweight, task-specific reasoning strategies can deliver strong performance on resource-constrained multimodal question answering benchmarks, without requiring massive compute budgets.
---
*Auto-collected on 2026-07-14*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178395120