Skip to content
Hot eventLive

社交媒体情感分析中人类、专用工具与LLM的一致性研究

1 reports1 sources4 hr ago updated

Get the story

AI overview

该研究评估 TextBlob、VADER、Twitter-roBERTa-base 与 Qwen-32B、GPT-OSS-120B、Llama-4-Maverick-17B 在 100 条推文上与 6 名人类标注者的一致性。研究使用 Cohen's kappa 与 Fleiss' kappa 作为衡量指标,结果显示人类标注者之间仅达到 fair agreement,表明在社交媒体文本情感分析任务上,人类、专用工具与 LLM 均难以形成共识。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Computation and Language
    人类、专用工具与 LLM 在社交媒体文本情感分析上均难达共识

    研究评估 TextBlob、VADER、Twitter-roBERTa-base 与 Qwen3-32B、GPT-OSS-120B、Llama-4-Maverick-17B 在 100 条推文上与 6 名人类标注者的一致性,Cohen's kappa 与 Fleiss' kappa 显示人类间仅 fair agreement。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.