[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-societybench-social-event-forecasting-benchmark-en":3,"article-related-societybench-social-event-forecasting-benchmark-en":29,"series-research-7591c5c2-467f-4859-9014-43f7c22bc136":75},{"id":4,"slug":5,"title":6,"content":7,"summary":8,"source":9,"source_url":10,"author":11,"image_url":12,"cover_image":12,"category":13,"language":14,"translated_content":11,"related_article_id":15,"keywords":16,"key_takeaways":22,"views":26,"created_at":27,"published_at":28,"topic_cluster_id":11},"7591c5c2-467f-4859-9014-43f7c22bc136","societybench-social-event-forecasting-benchmark-en","SocietyBench tests social-event forecasting","\u003Cp data-speakable=\"summary\">SocietyBench measures whether \u003Ca href=\"\u002Ftag\u002Fllms\">LLMs\u003C\u002Fa> can forecast how real social events unfold, using anonymized counterfactual timelines.\u003C\u002Fp>\u003Cul>\u003Cli>\u003Cstrong>Research org\u003C\u002Fstrong>: Unspecified in arXiv abstract\u003C\u002Fli>\u003Cli>\u003Cstrong>Core data\u003C\u002Fstrong>: 75.0\u002F100 on five events and 125 prediction points\u003C\u002Fli>\u003Cli>\u003Cstrong>Breakthrough\u003C\u002Fstrong>: Counterfactual timelines with entity replacement and date shifting\u003C\u002Fli>\u003C\u002Ful>\u003Cp>Most \u003Ca href=\"\u002Ftag\u002Fllm\">LLM\u003C\u002Fa> evaluation still asks a narrow question: can the model finish a task? This paper argues that leaves out a different capability that matters in the real world — whether a model can understand and forecast how social events evolve over time.\u003C\u002Fp>\u003Cp>That gap matters for anyone building agents that need to reason about news, public opinion, or event trajectories. If a system can browse, summarize, or classify but cannot track how an event changes across days and platforms, it may look competent while still missing the kind of temporal reasoning that real deployments need.\u003C\u002Fp>\u003Ch2>What problem SocietyBench is trying to fix\u003C\u002Fh2>\u003Cp>The authors say current benchmarks heavily measure task completion: fix a bug, drive a browser, operate a GUI. Those are useful, but they do not test a complementary social ability — forecasting how real social events unfold.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785911578381-50xp.png\" alt=\"SocietyBench tests social-event forecasting\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>SocietyBench is designed to measure that ability end to end. It starts from a one-line event topic, gathers web news and social-media posts from five platforms, turns them into a date-indexed timeline, and separates factual events from a public-opinion layer. Then every cutoff date on that timeline becomes a set of forecasting questions.\u003C\u002Fp>\u003Cp>The key idea is that this is not just a retrieval \u003Ca href=\"\u002Ftag\u002Fbenchmark\">benchmark\u003C\u002Fa> or a static QA set. It is built around evolution over time, which means a model has to reason about what was known by a given date and what likely comes next.\u003C\u002Fp>\u003Ch2>How the method works in plain English\u003C\u002Fh2>\u003Cp>SocietyBench builds a structured timeline first, then asks forecasting questions against that timeline. That timeline keeps two layers apart: the factual sequence of events and the public-opinion signals around them.\u003C\u002Fp>\u003Cp>Before any model sees the data, the benchmark applies a three-phase anonymization procedure. It replaces named entities and shifts every date by a per-event constant. The goal is to turn a real event arc into a counterfactual social world that is structurally the same, but stripped of labels a model might memorize from pre-training.\u003C\u002Fp>\u003Cp>In other words, the benchmark tries to reduce shortcut learning. A model should not be able to win just because it recognizes a famous event name or remembers the exact dates from training data.\u003C\u002Fp>\u003Cp>Each forecasting question is scored on two separate 100-point axes: probability calibration and temporal accuracy. That split is important because the paper treats “how likely” and “when” as different \u003Ca href=\"\u002Ftag\u002Fskills\">skills\u003C\u002Fa>, not one blended score.\u003C\u002Fp>\u003Ch2>What the paper actually shows\u003C\u002Fh2>\u003Cp>The evaluation covers five heterogeneous events and 125 prediction points in Chinese and English editions. The strongest of six frontier LLMs reaches 75.0 out of 100, compared with a trivial anchor of 50.\u003C\u002Fp>\n\u003Cfigure class=\"my-6\">\u003Cimg src=\"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785911576128-lbh0.png\" alt=\"SocietyBench tests social-event forecasting\" class=\"rounded-xl w-full\" loading=\"lazy\" \u002F>\u003C\u002Ffigure>\n\u003Cp>The paper says the two score axes can diverge. A model may be strong on calibration but weak on time, or the reverse. That means a single overall score can hide real failure modes.\u003C\u002Fp>\u003Cp>The authors also test three \u003Ca href=\"\u002Ftag\u002Fagent\">agent\u003C\u002Fa> frameworks built on a shared base model. None of them improve on that base model. Two model-free heuristics perform worse than every LLM tested.\u003C\u002Fp>\u003Cp>One of the clearest takeaways is that performance varies a lot by event. The paper reports per-event gaps as large as 21.4 points on a single axis, which is why the authors argue for evaluating across multiple events instead of relying on one.\u003C\u002Fp>\u003Cp>The abstract does not provide a full benchmark table, per-model breakdown, or detailed ablation numbers beyond those headline results. So the safe read is that SocietyBench exposes a difficult forecasting problem, not that one specific model has solved it.\u003C\u002Fp>\u003Ch2>Why developers should care\u003C\u002Fh2>\u003Cp>If you are building agents that ingest news, monitor social platforms, or track public narratives, this benchmark is a reminder that “knowing what happened” is not the same as “forecasting what happens next.”\u003C\u002Fp>\u003Cp>That distinction matters for product design. A system that can summarize a timeline but not separate factual events from opinion signals may be brittle in workflows like risk monitoring, trend analysis, or event-driven alerting.\u003C\u002Fp>\u003Cp>It also matters for evaluation. SocietyBench’s counterfactual setup is meant to make memorization harder, which is useful if you want to know whether your model is reasoning over structure instead of recalling a familiar story.\u003C\u002Fp>\u003Ch2>Limitations and open questions\u003C\u002Fh2>\u003Cp>The benchmark is promising, but the abstract leaves some important questions open. We do not get the full event list, the exact five platforms, the scoring rubric details, or the kinds of forecasting questions used for each cutoff date.\u003C\u002Fp>\u003Cp>We also do not know from the abstract how robust the anonymization is in practice, or whether some residual clues still leak the original event identity. That matters because counterfactual design only works if the shortcut removal is strong enough.\u003C\u002Fp>\u003Cp>Another open question is generality. The paper shows results on five heterogeneous events, but that is still a small slice of the social world. The authors themselves point to this by emphasizing multi-event evaluation rather than a single benchmark instance.\u003C\u002Fp>\u003Cp>Finally, the results suggest that adding agent scaffolding is not automatically enough. Three agent frameworks failed to beat the shared base model, which is a useful warning for teams assuming orchestration alone will solve social forecasting.\u003C\u002Fp>\u003Ch2>Bottom line\u003C\u002Fh2>\u003Cp>SocietyBench gives the field a way to test whether LLMs can forecast event evolution, not just answer questions about it. For engineers, the practical value is clear: it targets a real failure mode in news-aware and socially aware systems, and it does so with a counterfactual design that makes memorization harder.\u003C\u002Fp>\u003Cp>The paper’s main contribution is the benchmark itself: anonymized timelines, question banks, ground truth, and scoring code are all released. The main result is more sobering — even frontier models still leave room for improvement on this kind of temporal social reasoning.\u003C\u002Fp>","SocietyBench measures whether LLMs can forecast how real social events unfold, using anonymized counterfactual timelines.","arxiv.org","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.04009",null,"https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785911578381-50xp.png","research","en","fa03dc7f-4db2-4122-ab50-729e2f795964",[17,18,19,20,21],"LLM evaluation","social forecasting","counterfactual benchmark","temporal reasoning","agents",[23,24,25],"SocietyBench tests forecasting of social-event evolution, not just task completion.","It uses anonymized counterfactual timelines to reduce memorization shortcuts.","Frontier models still show uneven performance across calibration and time.",1,"2026-08-05T06:32:29.805432+00:00","2026-08-05T06:32:29.793+00:00",{"tags":30,"relatedLang":34,"relatedPosts":38},[31,33],{"name":17,"slug":32},"llm-evaluation",{"name":21,"slug":21},{"id":15,"slug":35,"title":36,"language":37},"societybench-social-event-forecasting-benchmark-zh","SocietyBench：測 LLM 社會事件預測","zh",[39,45,51,57,63,69],{"id":40,"slug":41,"title":42,"cover_image":43,"image_url":43,"created_at":44,"category":13},"cc6ec2ef-409d-41f4-8f9e-061ac1b580a5","anthropic-security-evals-real-internet-failure-en","Anthropic’s security evals are failing on the real internet","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785931381074-zsmb.png","2026-08-05T12:02:34.38877+00:00",{"id":46,"slug":47,"title":48,"cover_image":49,"image_url":49,"created_at":50,"category":13},"c3d7a875-9f48-40e7-9ab3-d3065c677f26","worldcup-arena-live-llm-forecasting-en","WorldCup Arena Tests LLM Forecasting Live","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785913375732-wutc.png","2026-08-05T07:02:29.588506+00:00",{"id":52,"slug":53,"title":54,"cover_image":55,"image_url":55,"created_at":56,"category":13},"609c0bbc-21fa-4cdc-9836-149e0a140201","parvl-parallel-scaling-multimodal-llms-en","ParVL scales multimodal LLMs in parallel","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785909771185-6vxr.png","2026-08-05T06:02:27.507133+00:00",{"id":58,"slug":59,"title":60,"cover_image":61,"image_url":61,"created_at":62,"category":13},"926bc32a-f0f6-4f54-8c67-437870ebc62c","onepot-bench-0-lab-aware-chemistry-benchmarks-en","onepot-Bench 0 tests lab-aware chemistry models","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785826984023-8yss.png","2026-08-04T07:02:34.694503+00:00",{"id":64,"slug":65,"title":66,"cover_image":67,"image_url":67,"created_at":68,"category":13},"e4e66f1a-2c10-4c30-8c5c-fe98e008d637","aurora-lm-continuous-latent-diffusion-text-en","AURORA-LM brings diffusion to text latents","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785823376131-zpdd.png","2026-08-04T06:02:30.975616+00:00",{"id":70,"slug":71,"title":72,"cover_image":73,"image_url":73,"created_at":74,"category":13},"12c649a1-45b8-4823-b59a-f9ce8a52c9fb","kimi-k3-is-already-doing-its-own-job-en","Kimi K3 Is Already Doing Its Own Job","https:\u002F\u002Fxxdpdyhzhpamafnrdkyq.supabase.co\u002Fstorage\u002Fv1\u002Fobject\u002Fpublic\u002Fcovers\u002Finline-1785808979724-vm3a.png","2026-08-04T02:02:35.463745+00:00",[76,81,86,91,96,101,106,111,116,121],{"id":77,"slug":78,"title":79,"created_at":80},"a2715e72-1fe8-41b3-abb1-d0cf1f710189","ai-predictions-2026-big-changes-en","AI Predictions for 2026: Brace for Big Changes","2026-03-26T01:25:07.788356+00:00",{"id":82,"slug":83,"title":84,"created_at":85},"8404bd7b-4c2f-4109-9ec4-baf29d88af2b","ml-papers-of-the-week-github-research-desk-en","ML Papers of the Week Turns GitHub Into a Research Desk","2026-03-27T01:11:39.480259+00:00",{"id":87,"slug":88,"title":89,"created_at":90},"87897a94-8065-4464-a016-1f23e89e17cc","ai-ml-conferences-to-watch-in-2026-en","AI\u002FML Conferences to Watch in 2026","2026-03-27T01:51:54.184108+00:00",{"id":92,"slug":93,"title":94,"created_at":95},"6f1987cf-25f3-47a4-b3e6-db0997695be8","openclaw-agents-manipulated-self-sabotage-en","OpenClaw Agents Can Be Manipulated Into Failure","2026-03-28T03:03:18.899465+00:00",{"id":97,"slug":98,"title":99,"created_at":100},"a53571ad-735a-4178-9f93-cb09b699d99c","vega-driving-language-instructions-en","Vega: Driving with Natural Language Instructions","2026-03-28T14:54:04.698882+00:00",{"id":102,"slug":103,"title":104,"created_at":105},"a34581d6-f36e-46da-88bb-582fb3e7425c","personalizing-autonomous-driving-styles-en","Drive My Way: Personalizing Autonomous Driving Styles","2026-03-28T14:54:26.148181+00:00",{"id":107,"slug":108,"title":109,"created_at":110},"2bc1ad7f-26ce-4f02-9885-803b35fd229d","training-knowledge-bases-writeback-rag-en","Training Knowledge Bases with WriteBack-RAG","2026-03-28T14:54:45.643433+00:00",{"id":112,"slug":113,"title":114,"created_at":115},"71adc507-3c54-4605-bbe2-c966acd6187e","packforcing-long-video-generation-en","PackForcing: Efficient Long-Video Generation Method","2026-03-28T14:55:02.646943+00:00",{"id":117,"slug":118,"title":119,"created_at":120},"675942ef-b9ec-4c5f-a997-381250b6eacb","pixelsmile-facial-expression-editing-en","PixelSmile Framework Enhances Facial Expression Editing","2026-03-28T14:55:20.633463+00:00",{"id":122,"slug":123,"title":124,"created_at":125},"6954fa2b-8b66-4839-884b-e46f89fa1bc3","adaptive-block-scaled-data-types-en","IF4: Smarter 4-Bit Quantization That Adapts to Your Data","2026-03-31T06:00:36.65963+00:00"]