{
  "title": "Daily Research Brief 2026-08-22",
  "url": "/en/posts/research-brief-2026-08-22/",
  "permalink": "https://hackcv.com/en/posts/research-brief-2026-08-22/",
  "date": "2026-08-22",
  "lastmod": "2026-08-22",
  "author": "",
  "description": "Daily research brief — AI / LLM / Agent / Computer Vision / Audio-Video / Engineering",
  "categories": ["Research Brief"],
  "tags": ["AI","LLM","Agent","Computer Vision","Audio-Video","Engineering","Daily Brief"],
  "cover": "https://picsum.photos/seed/daily-research-brief-2026-08-22/1200/675",
  "readingTime": 2,
  "wordCount": 396,
  "content": "\u003ch1 id=\"daily-research-brief-2026-08-22\"\u003eDaily Research Brief 2026-08-22\u003c/h1\u003e\n\u003cp\u003e📊 Token usage: estimated from retrieval and writing scale.\u003c/p\u003e\n\u003cp\u003eCovers the latest AI research, open source and industry moves, updated daily.\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"editors-note\"\u003eEditor\u0026rsquo;s Note\u003c/h2\u003e\n\u003cp\u003eToday\u0026rsquo;s main thread: \u0026ldquo;execution systems + skill ecosystems\u0026rdquo; formally take over the leverage point of AI competition, with long-context inference efficiency and agent memory as two technical undercurrents.\u003c/p\u003e\n\u003ch2 id=\"1-latest-arxiv-papers\"\u003e1. Latest arXiv Papers\u003c/h2\u003e\n\u003ch3 id=\"1-envharness-awakening-static-worlds-for-agent-learning\"\u003e1. EnvHarness: Awakening Static Worlds for Agent Learning\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003eAbstract\u003c/strong\u003e: A framework that turns static repositories into dynamic, evolving environments for agent RL — no domain-specific customization or expensive verifiers needed. Environments co-evolve with the policy during training.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLink\u003c/strong\u003e: \u003ca href=\"https://arxiv.org/abs/2608.19880\"\u003ehttps://arxiv.org/abs/2608.19880\u003c/a\u003e\u003c/p\u003e\n\u003ch3 id=\"2-vla-self-demo-fine-tuning\"\u003e2. VLA Self-Demo Fine-Tuning\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003eAbstract\u003c/strong\u003e: Vision-language-action models fine-tuned on self-generated demonstrations for long-horizon manipulation (+11.6%), zero parameter updates to the base policy.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLink\u003c/strong\u003e: \u003ca href=\"https://arxiv.org/abs/2608.19490\"\u003ehttps://arxiv.org/abs/2608.19490\u003c/a\u003e\u003c/p\u003e\n\u003ch3 id=\"3-flashprefill-v2-block-sparse-prefill-attention\"\u003e3. FlashPrefill V2: Block-Sparse Prefill Attention\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003eAbstract\u003c/strong\u003e: Block-sparse prefill attention that cuts KV and attention compute for long-context prefill.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLink\u003c/strong\u003e: \u003ca href=\"https://arxiv.org/abs/2608.19758\"\u003ehttps://arxiv.org/abs/2608.19758\u003c/a\u003e\u003c/p\u003e\n\u003ch3 id=\"4-swe-bench-science\"\u003e4. SWE-bench Science\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003eAbstract\u003c/strong\u003e: A scientific-reproduction variant of SWE-bench evaluating agents on faithfully reproducing papers.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLink\u003c/strong\u003e: \u003ca href=\"https://arxiv.org/abs/2608.19799\"\u003ehttps://arxiv.org/abs/2608.19799\u003c/a\u003e\u003c/p\u003e\n\u003ch3 id=\"5-personalbench-what-personalized-llms-reveal-about-author-identity\"\u003e5. PersonalBench: What Personalized LLMs Reveal About Author Identity\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003eAbstract\u003c/strong\u003e: A benchmark probing how personalized LLMs reflect author identity — and what that reveals about attribution.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLink\u003c/strong\u003e: \u003ca href=\"https://arxiv.org/abs/2608.19746\"\u003ehttps://arxiv.org/abs/2608.19746\u003c/a\u003e\u003c/p\u003e\n\u003ch3 id=\"6-recache-tool-augmented-agent-kv-cache-reuse\"\u003e6. ReCache: Tool-Augmented Agent KV-Cache Reuse\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003eAbstract\u003c/strong\u003e: KV-cache reuse across tool-augmented agent steps — cutting redundant recomputation in long tool loops.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLink\u003c/strong\u003e: \u003ca href=\"https://arxiv.org/abs/2608.19662\"\u003ehttps://arxiv.org/abs/2608.19662\u003c/a\u003e\u003c/p\u003e\n\u003ch3 id=\"7-can-agent-memory-systems-track-evolving-state\"\u003e7. Can Agent Memory Systems Track Evolving State?\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003eAbstract\u003c/strong\u003e: Evaluates agent memory systems on tracking evolving state (StateMemBench), showing current-state accuracy lifts from 0.205 to 0.363 on DeepSeek-V4-Flash (1.8×).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLink\u003c/strong\u003e: \u003ca href=\"https://arxiv.org/abs/2608.19652\"\u003ehttps://arxiv.org/abs/2608.19652\u003c/a\u003e\u003c/p\u003e\n\u003ch3 id=\"8-eureka-task-conditioned-meta-agent-orchestration-for-scientific-discovery\"\u003e8. Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003eAbstract\u003c/strong\u003e: Task-conditioned meta-agent orchestration for autonomous scientific discovery.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLink\u003c/strong\u003e: \u003ca href=\"https://arxiv.org/abs/2608.19047\"\u003ehttps://arxiv.org/abs/2608.19047\u003c/a\u003e\u003c/p\u003e\n\u003ch2 id=\"2-hot-github-open-source\"\u003e2. Hot GitHub Open Source\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eDeepSeek deepseek-harness\u003c/strong\u003e — \u0026ldquo;everything is a plugin\u0026rdquo; harness, 130k★ in 4 days\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003ererelease of OpenAI Codex runtime\u003c/strong\u003e (Apache-2.0) — reasoning-trace retention + context compression lifted ARC-AGI-3 from 13.3% to 38.3% with 1/6 output tokens\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAgent skills wave\u003c/strong\u003e: addyosmani/agent-skills (80k★), obra/superpowers, pbakaus/impeccable, book-to-skill, spec-kit, headroom (context compression, 60–95% token cut)\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"3-selected-industry-news\"\u003e3. Selected Industry News\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eAnthropic\u003c/strong\u003e: Computer Use / Browser Use / Skills API / Files API all GA\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eNVIDIA AVO\u003c/strong\u003e: same Claude Opus 5 hit 25/25 (100 RHAE) on ARC-AGI-3 public set; 7 days of GPU kernels up to 3.5% faster than cuDNN\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003ePricing\u003c/strong\u003e: DeepSeek weekend valley pricing; OpenAI GPT-5.6 Sol −33% output price ($30→$20); Gemini 3.7 Flash ~half price\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eSecurity\u003c/strong\u003e: OpenAI pauses two weeks of large-scale training after HF breach; Anthropic archives \u0026ldquo;Model 2\u0026rdquo;; China\u0026rsquo;s mandatory agent-security national standard moves forward\u003c/li\u003e\n\u003c/ul\u003e\n",
  "summary": "Daily Research Brief 2026-08-22 📊 Token usage: estimated from retrieval and writing scale.\nCovers the latest AI research, open source and industry moves, updated daily.\nEditor\u0026rsquo;s Note Today\u0026rsquo;s main thread: \u0026ldquo;execution systems + skill ecosystems\u0026rdquo; formally take over the leverage point of AI competition, with long-context inference efficiency and agent memory as two technical undercurrents.\n1. Latest arXiv Papers 1. EnvHarness: Awakening Static Worlds for Agent Learning Abstract: A framework that turns static repositories into dynamic, evolving environments for agent RL — no domain-specific customization or expensive verifiers needed. Environments co-evolve with the policy during training.\n"
}
