Overview
Fields: cs.AI, cs.CL, cs.SI Authors: AS Aravinthakshan, Laven Srivastava, Harsh Nandwani Published: 2026-09-13 arXiv: 2609.11144
Abstract (full translation)
Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.
Key findings
- Human-label validation (construct validity) and market-signal extraction (predictive validity) are not interchangeable measures of tool quality.
- Results depend on sampling conventions: method-specific sampling favors same-day associations; a fixed-n panel yields similar rank correlations at both same-day and one-day-ahead horizons.
- Coarse (binary) score ordering correlates weakly with predictive performance regardless of horizon.
- Message volume does not predict market damage or settlement size in conversations where 17.6% of messages are spam.