{
  "title": "Daily Research Brief 2026-08-20",
  "url": "/en/posts/research-brief-2026-08-20/",
  "permalink": "https://hackcv.com/en/posts/research-brief-2026-08-20/",
  "date": "2026-08-20",
  "lastmod": "2026-08-20",
  "author": "",
  "description": "Daily research brief — AI / LLM / Agent / Computer Vision / Audio-Video / Engineering",
  "categories": ["Research Brief"],
  "tags": ["AI","LLM","Agent","Computer Vision","Audio-Video","Engineering","Daily Brief"],
  "cover": "https://picsum.photos/seed/daily-research-brief-2026-08-20/1200/675",
  "readingTime": 2,
  "wordCount": 326,
  "content": "\u003ch1 id=\"daily-research-brief-2026-08-20\"\u003eDaily Research Brief 2026-08-20\u003c/h1\u003e\n\u003cp\u003e📊 Token usage: estimated from retrieval and writing scale.\u003c/p\u003e\n\u003cp\u003eCovers the latest AI research, open source and industry moves, updated daily.\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"editors-note\"\u003eEditor\u0026rsquo;s Note\u003c/h2\u003e\n\u003cp\u003eThe one thing worth recording today: capability increments are moving from \u0026ldquo;model weights\u0026rdquo; to \u0026ldquo;execution systems + authorization boundaries\u0026rdquo;. StateM spent nothing on training — only rebuilding the harness (persistent state, staged context, verifiable transitions, recoverable runbooks) — to push Terminal-Bench 2.1 raw accuracy to 95.3%, or hit the same score at ~$15 of API spend vs the GPT reference line\u0026rsquo;s $574.68. The same day, Demystifying Agent Skills used 8,135 trial records to explain why skills work: \u003cstrong\u003e65.7% of gains come from \u0026ldquo;program anchoring\u0026rdquo;, not injected knowledge\u003c/strong\u003e, and retrieval precision collapses from 29.6% to 3.3% as the skill pool grows from 5 to 100.\u003c/p\u003e\n\u003ch2 id=\"1-latest-arxiv-papers\"\u003e1. Latest arXiv Papers\u003c/h2\u003e\n\u003col\u003e\n\u003cli\u003e\u003cstrong\u003eWhat is Missing from AI Post-Training AI: An Empirical Analysis\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2608.19072\"\u003ehttps://arxiv.org/abs/2608.19072\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eBayesian Partner Modelling enables Adaptive Replanning for LLM Coordination\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2608.18490\"\u003ehttps://arxiv.org/abs/2608.18490\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eStateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2608.15089\"\u003ehttps://arxiv.org/abs/2608.15089\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eDemystifying Agent Skills: Why They Work-Until They Don\u0026rsquo;t\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2608.14036\"\u003ehttps://arxiv.org/abs/2608.14036\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eASI-Bench: At the Dawn of Artificial Superintelligence\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2608.17271\"\u003ehttps://arxiv.org/abs/2608.17271\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eWhen Agents Act on Web3: An Attack-Surface Survey of MCP, Skills, and Tool Calling\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2608.17275\"\u003ehttps://arxiv.org/abs/2608.17275\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003ePhysics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2608.16578\"\u003ehttps://arxiv.org/abs/2608.16578\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eEmbodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2608.17512\"\u003ehttps://arxiv.org/abs/2608.17512\u003c/a\u003e\u003c/li\u003e\n\u003c/ol\u003e\n\u003ch2 id=\"2-hot-github-open-source\"\u003e2. Hot GitHub Open Source\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003edeepseek-harness\u003c/strong\u003e ecosystem surging (130k★ in 4 days); task-aware routing suites (dsh-routing-suite, sprix-sage-router) charting independently\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eponytail\u003c/strong\u003e (111.8k★) — \u0026ldquo;cognitive restraint, default-don\u0026rsquo;t-implement\u0026rdquo; agents\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eOpenBot\u003c/strong\u003e (CopilotKit) — containerized agents with review-before-act governance gates\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"3-selected-industry-news\"\u003e3. Selected Industry News\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eStateM vs GPT reference line\u003c/strong\u003e: harness-only scaling reaches frontier results at ~$15 vs $574.68\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eOpenAI\u003c/strong\u003e: admits underestimating model offensive cyber capability (HF incident); pauses two weeks of large-scale training\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAnthropic\u003c/strong\u003e: watermarks on all models; archives \u0026ldquo;Model 2\u0026rdquo; over alignment risk\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eNVIDIA AVO\u003c/strong\u003e: same Claude Opus 5 hit 25/25 on ARC-AGI-3 public set\u003c/li\u003e\n\u003c/ul\u003e\n",
  "summary": "Daily Research Brief 2026-08-20 📊 Token usage: estimated from retrieval and writing scale.\nCovers the latest AI research, open source and industry moves, updated daily.\nEditor\u0026rsquo;s Note The one thing worth recording today: capability increments are moving from \u0026ldquo;model weights\u0026rdquo; to \u0026ldquo;execution systems + authorization boundaries\u0026rdquo;. StateM spent nothing on training — only rebuilding the harness (persistent state, staged context, verifiable transitions, recoverable runbooks) — to push Terminal-Bench 2.1 raw accuracy to 95.3%, or hit the same score at ~$15 of API spend vs the GPT reference line\u0026rsquo;s $574.68. The same day, Demystifying Agent Skills used 8,135 trial records to explain why skills work: 65.7% of gains come from \u0026ldquo;program anchoring\u0026rdquo;, not injected knowledge, and retrieval precision collapses from 29.6% to 3.3% as the skill pool grows from 5 to 100.\n"
}
