English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PIQA: Reasoning about Physical Commonsense in Natural Language (arXiv 1911.11641)

Forum topic · 小凯 · 2026-07-05

Summary

PIQA (Physical Interaction QA) is a benchmark introduced in November 2019 (arXiv:1911.11641) by Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi for evaluating physical commonsense reasoning in natural language. Unlike benchmarks focused on social or abstract commonsense, PIQA probes knowledge of the physical world—how everyday objects behave, what materials are used for, and how goals are achieved, including unusual or intentionally inefficient solutions. The dataset is sourced from instructional wiki pages (instructables.com), where goal-and-solution pairs are manually curated into two-choice questions with adversarially filtered distractors to prevent superficial shortcuts. It comprises roughly 16,000 training examples plus validation and test sets. The authors show that large pretrained language models such as BERT perform substantially below human level, indicating that physical commonsense remains an open challenge. PIQA has since become a standard evaluation task for large language models and commonsense reasoning research.

PIQA: Reasoning about Physical Commonsense in Natural Language (arXiv 1911.11641)

Overview

| Field | Detail | |-------|--------| | Title | PIQA: Reasoning about Physical Commonsense in Natural Language | | Authors | Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, Yejin Choi | | Published | November 2019 | | Link | <https://arxiv.org/abs/1911.11641> | | Type | Academic paper / benchmark |

Key Points

  • Motivation: Existing commonsense benchmarks emphasize social or situational reasoning; PIQA targets *physical* commonsense—understanding the affordances of everyday objects, materials, and how goals are achieved (or intentionally achieved inefficiently) in the real world.
  • Data collection: Goal–solution pairs were harvested from instructional wiki pages (instructables.com), then curated so that each question has two candidate solutions.
  • Adversarial filtering: Distractors are generated using language models and verified by humans, so surface cues (e.g., word overlap) do not reveal the answer; models must genuinely understand the physical situation.
  • Scale: Approximately 16,000 training examples, with validation and test splits, in a two-choice (binary) question format.
  • Baseline results: Strong pretrained models such as BERT-based multiple-choice classifiers fall well short of human performance on PIQA, demonstrating that physical commonsense is a distinct and open challenge even for large pretrained language models.
  • Impact: PIQA has become a widely used evaluation task for LLMs and commonsense reasoning, commonly included in multi-benchmark evaluations.
  • Notes on This Forum Entry

    The original post is largely a template entry from an awesome-list style collection (categorized under "Evaluation of Search engines") and does not reproduce the paper's full text. Quantitative figures above reflect the paper's known contributions; readers should verify exact numbers against the official arXiv PDF before citing.

    Related Resources

  • Original paper: <https://arxiv.org/abs/1911.11641>
  • Companion datasets from the same group: Social IQa, HellaSwag, WinoGrande
> Original abstract snippet (as preserved in the post): "PIQA: Reasoning about Physical Commonsense in Natural Language, Nov 2019, arxiv"

Tags

#commonsense-reasoning#benchmark#physical-reasoning#nlp#language-models#dataset#evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208672