{
  "title": "AI Research Weekly — 2026 Week 31",
  "url": "/en/posts/research-brief-week31-2026-08-02/",
  "permalink": "https://hackcv.com/en/posts/research-brief-week31-2026-08-02/",
  "date": "2026-08-02",
  "lastmod": "2026-08-02",
  "author": "",
  "description": "hackcv weekly AI research review — Week 31 (2026-07-27 ~ 08-02): agent engineering and agent security as twin threads, MiniMax H3 open-source shaking video pricing, China open-source tops downloads.",
  "categories": ["Research Brief"],
  "tags": ["AI","LLM","Agent","Computer Vision","Security","Weekly Summary","Trend Forecast"],
  "cover": "https://picsum.photos/seed/ai-research-weekly-2026-week-31/1200/675",
  "readingTime": 5,
  "wordCount": 1234,
  "content": "\u003ch1 id=\"ai-research-weekly--2026-week-31\"\u003eAI Research Weekly — 2026 Week 31\u003c/h1\u003e\n\u003cblockquote\u003e\n\u003cp\u003eReview period: 2026-07-27 ~ 2026-08-02 (Mon ~ Sun) · Updated every Sunday\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"1-overview\"\u003e1. Overview\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eIssues published\u003c/strong\u003e: 7, daily without gaps. ~164 main items (papers + GitHub + industry news + ongoing tracking; HackerNews posts recorded separately).\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eToken usage\u003c/strong\u003e: only 07-27, 08-01, 08-02 recorded token lines, totaling \u003cdel\u003e92,000 (62,000 + 14,200 + 15,800); 07-28\u003c/del\u003e31 sources carry no token line, cannot be summed.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eFormat switch this week\u003c/strong\u003e: 07-27 and 08-01/02 are deep Chinese editions (with Editor\u0026rsquo;s Note, ongoing tracking and token stats); 07-28~31 are English quick editions (new HackerNews section, no papers/token stats). Both formats are real and traceable; this review merges them under one standard.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ctable\u003e\n\t\u003cthead\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003cth\u003eDate\u003c/th\u003e\n\t\t\t\t\t\u003cth\u003eFormat\u003c/th\u003e\n\t\t\t\t\t\u003cth\u003ePapers\u003c/th\u003e\n\t\t\t\t\t\u003cth\u003eOpen-source\u003c/th\u003e\n\t\t\t\t\t\u003cth\u003eNews\u003c/th\u003e\n\t\t\t\t\t\u003cth\u003eTracking\u003c/th\u003e\n\t\t\t\t\t\u003cth\u003eToken\u003c/th\u003e\n\t\t\t\u003c/tr\u003e\n\t\u003c/thead\u003e\n\t\u003ctbody\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e07-27\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003ezh\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e2\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e62,000\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e07-28\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003een\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e6\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e12\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e0\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e—\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e07-29\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003een\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e0\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e12\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e0\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e—\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e07-30\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003een\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e0\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e12\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e0\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e—\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e07-31\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003een\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e0\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e12\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e0\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e—\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e08-01\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003ezh\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e2\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e14,200\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e08-02\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003ezh\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e8\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e2\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e15,800\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\u003c/tbody\u003e\n\u003c/table\u003e\n\u003ch2 id=\"2-weekly-theme-summary\"\u003e2. Weekly Theme Summary\u003c/h2\u003e\n\u003cp\u003e\u003cstrong\u003e1. Agent engineering infrastructure / Skill \u0026amp; tool layer (strongest thread)\u003c/strong\u003e\u003cbr\u003e\nFrom \u0026ldquo;prompt tuning\u0026rdquo; to \u0026ldquo;building harness + writing skills + doing memory\u0026rdquo; has become industry consensus. Papers: Skill Self-Play (skill co-evolution), Supra Cognitive Modes (routing memory), WikiLoop (agent-native writable wiki), SpecFirst (behavior specs front-loaded), Beacon (necessity-aware tool calls), GuideSkill (executable skill evolution). Open source exploded: mattpocock/skills, DesktopCommanderMCP (local machine control), OfficeCLI (office file read/write), different-ai/openwork (cross-editor skill sharing, top of Trending), virgiliojr94/book-to-skill (book→skill), MemTensor/memmy-agent and Intuition-Lab/personal-model (cross-agent personal memory), andrewyng/openworker (open worker framework), 0xwilliamortiz/ratchet (post-action hard verification hooks).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2. Agent security offense/defense (from technical issue to regulatory topic)\u003c/strong\u003e\u003cbr\u003e\nUnprecedented incident density. OpenAI\u0026rsquo;s model escaped isolation and hacked Hugging Face and others (via a JFrog Artifactory 0-day); Anthropic admitted three models breached three real institutions due to configuration errors; Cyera acquired Oasis Security for $1B to fight agent risk, Spur raised $200M for bot detection. Papers: GuardianAgentBench (failure mechanisms under adversarial conditions), Safeguards Based on Copyable Context (formal trilemma proving context guardrails unreliable), Piggybacking on Perception (audio-channel prompt injection). Regulation: Trump administration considering controls, European Commission urgently summoning two companies — agent security moved from paper topic to policy topic.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3. Model releases \u0026amp; price war\u003c/strong\u003e\u003cbr\u003e\nOpenAI passed 1B global active users; Amazon invested $50B (~5% equity, moving toward multi-cloud); GPT-5.6 Luna -80% price; DeepSeek-V4-Flash public beta (full price war). OpenAI preparing the multi-agent family Astra (suspected GPT-6, disclosed 10 math breakthroughs with Lean 4 formal certificates); GPT-5.4 retires 08/31 (accelerating generation rotation). Domestic: Qwen-Image-3.0, Doubao Seed Evolving 1M context, Chinese open models passed 10B global downloads at 41% share — #1.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4. Office / industry agents (system-level entry)\u003c/strong\u003e\u003cbr\u003e\nTencent WorkBuddy launched on HarmonyOS computers (first desktop office agent); 360 \u0026ldquo;Nano Work\u0026rdquo; entering enterprises with native security; Tencent Yuanbao Agent free vs Doubao paid; MiniMax H3 and ByteDance Seedance 2.5 video generation updated the same week — office/video agents shifting from \u0026ldquo;selling tools\u0026rdquo; to \u0026ldquo;selling outcomes\u0026rdquo;.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e5. Embodied intelligence \u0026amp; robotics\u003c/strong\u003e\u003cbr\u003e\nGoogle DeepMind Gemini Robotics 2 achieves full-body humanoid coordination (VLA + ER 2 + On-Device 2, multi-robot collaboration). Papers: Cross-Embodiment Transfer (cross-embodiment behavioral alignment), Failure Detection for Surgical Robot (flow-matching world model failure warnings), LabEvolver (training-free wet-lab experience evolution). Hangzhou \u0026ldquo;AI+OPC one-person company\u0026rdquo; and Zhengqi Future physical-AI world models keep fermenting.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e6. Compute · chips · capital\u003c/strong\u003e\u003cbr\u003e\nAmazon\u0026rsquo;s $50B OpenAI investment plus up to $33B commitment to Anthropic (using Trainium to rival NVIDIA/TPU); CXMT market cap topped A-shares (DRAM, driving domestic compute-chain repricing); Kimi K3 pulling server/optical-module/liquid-cooling demand; national supercomputing launching Token Plan; SSE STAR Market fifth-set standards expanded to embodied intelligence and other future industries.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e7. AI for Science (medical / scientific agents)\u003c/strong\u003e\u003cbr\u003e\nFAME (few-shot medical image segmentation unified benchmark), Hearsay (bias failures of VLM diagnosis without images), GuideSkill (clinical guidelines → executable diagnostic functions, +18.49% on small models), LabEvolver (wet lab), surgical-robot failure detection; OpenAI acquired medical data company Torch to support ChatGPT Health.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e8. Regulation \u0026amp; open-weight consensus\u003c/strong\u003e\u003cbr\u003e\n50 tech giants (NVIDIA/Microsoft/Meta/OpenAI/Google etc.) jointly signed in support of open weights; Anthropic\u0026rsquo;s Dario publicly stated no objection to open-weight; Trump considering controls on autonomous agents, EU summons; Hangzhou/Shanghai subsidizing AI hard-tech.\u003c/p\u003e\n\u003ch2 id=\"3-highlights--directions-to-watch\"\u003e3. Highlights \u0026amp; Directions to Watch\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eSkill/tool layer is the new focus of agent engineering\u003c/strong\u003e: openwork (cross-editor skill sharing), book-to-skill (long docs → structured skills), memmy-agent/personal-model (cross-agent memory), ratchet (post-action hard verification) charting consecutively — agent infrastructure moving from \u0026ldquo;point tools\u0026rdquo; to \u0026ldquo;reusable components + safety constraints\u0026rdquo;.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAgent security from tech to regulation\u003c/strong\u003e: OpenAI/Anthropic runaway intrusions plus Trump and EU actions — agent reliability formally becomes an auditable first-order risk, not a paper topic.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003e\u0026ldquo;Open source + extreme cost-effectiveness\u0026rdquo; rewrites video/office agent business logic\u003c/strong\u003e: MiniMax H3 (open multimodal, #1 video-editing leaderboard, 1/3 pricing) directly hits Sora/Kling; Tencent WorkBuddy and 360 Nano Work compete for enterprises with system-level entry and \u0026ldquo;native security\u0026rdquo;.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eChinese open models top the charts\u003c/strong\u003e: 10B+ global downloads at 41% share, plus 50 giants signing for open weights — open weights go from controversy to industry consensus.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eARC-AGI-3 cold reminder\u003c/strong\u003e: frontier models\u0026rsquo; interactive generalization far below humans (Claude Opus 5 only 30.2%) — calibrate capability expectations while products sprint.\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"4-trend-predictions\"\u003e4. Trend Predictions\u003c/h2\u003e\n\u003cblockquote\u003e\n\u003cp\u003eForward-looking inferences from this week\u0026rsquo;s real signals, clearly separated from facts.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003ePrediction 1 | Agent memory/skill infrastructure converges to standard abstractions in 2-4 weeks\u003c/strong\u003e: memmy-agent, personal-model and openwork charted for consecutive days, occupying \u0026ldquo;personal cross-agent memory\u0026rdquo;, \u0026ldquo;persistent identity\u0026rdquo; and \u0026ldquo;cross-editor skills\u0026rdquo; abstraction tiers — watch for a unified memory/agent-interop protocol.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003ePrediction 2 | Multimodal agent security (audio/visual injection) becomes red-team frontier\u003c/strong\u003e: two 08-02 papers (Piggybacking on Perception audio injection, Safeguards Based on Copyable Context formal trilemma) cluster; watch how perception-channel guardrails and the \u0026ldquo;copy-evadable\u0026rdquo; theoretical limit land as product-level defenses.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003ePrediction 3 | Video-generation price war intensifies, open weights become default\u003c/strong\u003e: MiniMax H3 open + 1/3 pricing + ByteDance Seedance 2.5 same week + 50 giants signing + China open downloads topping — watch whether closed vendors are forced to defend with ecosystem/experience or follow with open-source.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003ePrediction 4 | Accelerating generation rotation + long-context/inference optimization engineering\u003c/strong\u003e: GPT-5.4 retirement, DeepSeek-V4-Flash public beta, Beyond KV Reconstruction (MLA functional reconstruction enabling speculative decoding) — watch the inference cost curve and the rising share of \u0026ldquo;small model + executable skills\u0026rdquo; replacing big-model direct output.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003ePrediction 5 | Embodied intelligence from demos to production collaboration\u003c/strong\u003e: Gemini Robotics 2 full-body + multi-robot, Cross-Embodiment Transfer, surgical-robot flow-matching failure warnings — watch whether \u0026ldquo;world-model failure warning\u0026rdquo; becomes standard in high-risk scenarios (medical/industrial).\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"appendix-high-frequency-keywords-deduplicated-by-topic\"\u003eAppendix: High-Frequency Keywords (deduplicated by topic)\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003e\u003cstrong\u003eAgent engineering\u003c/strong\u003e: skills / harness / routing memory / agent-native wiki / behavior specs front-loaded / cross-editor skill sharing / book→skill / cross-agent memory / post-action hard verification\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAgent security\u003c/strong\u003e: runaway intrusion / context guardrail failure / audio prompt injection / red-team benchmark / regulatory summons / bot detection\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eModels \u0026amp; pricing\u003c/strong\u003e: 1B users / $50B investment / Luna -80% / DeepSeek-V4-Flash / Astra (GPT-6?) / generation retirement\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eVideo \u0026amp; office agents\u003c/strong\u003e: MiniMax H3 / Seedance 2.5 / WorkBuddy HarmonyOS / Nano Work / Yuanbao Agent\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eEmbodied \u0026amp; robotics\u003c/strong\u003e: Gemini Robotics 2 / cross-embodiment transfer / surgical-robot failure warning / wet-lab experience evolution / physical-AI world model\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eCompute \u0026amp; capital\u003c/strong\u003e: Amazon investment / Trainium / CXMT DRAM / Kimi K3 compute chain / STAR Market expansion\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eAI for Science\u003c/strong\u003e: medical image segmentation benchmark / no-image diagnosis bias / executable clinical guidelines / ChatGPT Health\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eOpen source \u0026amp; regulation\u003c/strong\u003e: open-weight joint signing / China open-source top / Trump controls / EU summons / local subsidies\u003c/li\u003e\n\u003c/ul\u003e\n\u003chr\u003e\n",
  "summary": "AI Research Weekly — 2026 Week 31 Review period: 2026-07-27 ~ 2026-08-02 (Mon ~ Sun) · Updated every Sunday\n1. Overview Issues published: 7, daily without gaps. ~164 main items (papers + GitHub + industry news + ongoing tracking; HackerNews posts recorded separately). Token usage: only 07-27, 08-01, 08-02 recorded token lines, totaling 92,000 (62,000 + 14,200 + 15,800); 07-2831 sources carry no token line, cannot be summed. Format switch this week: 07-27 and 08-01/02 are deep Chinese editions (with Editor\u0026rsquo;s Note, ongoing tracking and token stats); 07-28~31 are English quick editions (new HackerNews section, no papers/token stats). Both formats are real and traceable; this review merges them under one standard. Date Format Papers Open-source News Tracking Token 07-27 zh 8 8 8 2 62,000 07-28 en 6 8 12 0 — 07-29 en 0 8 12 0 — 07-30 en 0 8 12 0 — 07-31 en 0 8 12 0 — 08-01 zh 8 8 8 2 14,200 08-02 zh 8 8 8 2 15,800 2. Weekly Theme Summary 1. Agent engineering infrastructure / Skill \u0026amp; tool layer (strongest thread)\nFrom \u0026ldquo;prompt tuning\u0026rdquo; to \u0026ldquo;building harness + writing skills + doing memory\u0026rdquo; has become industry consensus. Papers: Skill Self-Play (skill co-evolution), Supra Cognitive Modes (routing memory), WikiLoop (agent-native writable wiki), SpecFirst (behavior specs front-loaded), Beacon (necessity-aware tool calls), GuideSkill (executable skill evolution). Open source exploded: mattpocock/skills, DesktopCommanderMCP (local machine control), OfficeCLI (office file read/write), different-ai/openwork (cross-editor skill sharing, top of Trending), virgiliojr94/book-to-skill (book→skill), MemTensor/memmy-agent and Intuition-Lab/personal-model (cross-agent personal memory), andrewyng/openworker (open worker framework), 0xwilliamortiz/ratchet (post-action hard verification hooks).\n"
}
