论文概要
研究领域: cs.AI, cs.CL, cs.SI
作者: AS Aravinthkakshan, Laven Srivastava, Harsh Nandwani
发布时间: 2026-09-13
arXiv: 2609.11144
中文摘要
金融 NLP 有一个标准工作流程:针对人类标签验证情感工具,然后信任它提取市场信号。这假设两种评估测量相同的东西。我们在可以同时测量两者的设置中测试该假设:2002-2025 年证券集体诉讼语料库,将 70,500 条 X 消息与异常股票回报联系起来,使用单一注释者人类标记的黄金样本。通过同一管道运行五种工具(VADER、Loughran-McDonald、FinBERT、Twitter-RoBERTa 和 LLM 注释器),我们发现构造和预测有效性之间的关系取决于采样约定和分数表示。在常规方法特定采样下,人类一致性与分级同日关联更密切,而非一日领先。然而,在固定 n 面板上,一致性在两个时间范围具有相似的分级秩相关,而粗排序保持较弱。基准一致性因此建立语义有效性,但本身不决定预测排名。在 17.6% 垃圾邮件的对话中,消息量既无法预测市场损害也无法预测和解规模。
原文摘要
Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.
自动采集于 2026-09-13
#论文 #arXiv #AI #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。