{
  "title": "Daily Research Brief 2026-03-26",
  "url": "/en/posts/research-brief-2026-03-26/",
  "permalink": "https://hackcv.com/en/posts/research-brief-2026-03-26/",
  "date": "2026-03-26",
  "lastmod": "2026-03-26",
  "author": "",
  "description": "Daily research brief — AI / LLM / Agent / Computer Vision / Audio-Video / Engineering",
  "categories": ["Research Brief"],
  "tags": ["AI","LLM","Agent","Computer Vision","Audio-Video","Engineering","Daily Brief"],
  "cover": "https://picsum.photos/seed/daily-research-brief-2026-03-26/1200/675",
  "readingTime": 2,
  "wordCount": 424,
  "content": "\u003ch1 id=\"daily-research-brief-2026-03-26\"\u003eDaily Research Brief 2026-03-26\u003c/h1\u003e\n\u003cp\u003e📊 Token usage: estimated from retrieval and writing scale.\u003c/p\u003e\n\u003cp\u003eCovers the latest AI research, open source and industry moves, updated daily.\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"editors-note\"\u003eEditor\u0026rsquo;s Note\u003c/h2\u003e\n\u003cp\u003eToday\u0026rsquo;s paper thread: unified policy optimization for reasoning-driven visual generation (UniGRPO), generalized unconstrained urban 3D occupancy (OccAny), foveation-inspired efficient image/video generation, on-demand vision interaction for VLLM efficiency, and zero-shot referring video object segmentation (AgentRVOS). GitHub trending is dominated by agent infrastructure — Karpathy\u0026rsquo;s autoresearch, gstack, paperclip, CLI-Anything and Google Workspace CLI.\u003c/p\u003e\n\u003ch2 id=\"1-latest-arxiv-papers\"\u003e1. Latest arXiv Papers\u003c/h2\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eUniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation\u003c/strong\u003e — unifies text and image (autoregressive + flow matching) joint generation with a unified policy optimization method, UniGRPO, for reasoning-driven visual content generation. — \u003ca href=\"https://arxiv.org/abs/2603.23500\"\u003ehttps://arxiv.org/abs/2603.23500\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eOccAny: Generalized Unconstrained Urban 3D Occupancy\u003c/strong\u003e — breaks the dependence of 3D occupancy prediction on in-domain annotations and precise sensor calibration, proposing more generalizable unconstrained urban 3D occupancy prediction. — \u003ca href=\"https://arxiv.org/abs/2603.23502\"\u003ehttps://arxiv.org/abs/2603.23502\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eFoveated Diffusion: Efficient Spatially Adaptive Image and Video Generation\u003c/strong\u003e — borrows the foveal vision mechanism of the human eye for spatially adaptive, efficient diffusion/flow-matching image and video generation, significantly cutting compute. — \u003ca href=\"https://arxiv.org/abs/2603.23491\"\u003ehttps://arxiv.org/abs/2603.23491\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eVision On Request: Enhanced VLLM Efficiency with Sparse Dynamic Vision-Language Interactions\u003c/strong\u003e — on-demand vision interaction replaces conventional visual token pruning, greatly improving LVLM inference efficiency while preserving information fidelity. — \u003ca href=\"https://arxiv.org/abs/2603.23495\"\u003ehttps://arxiv.org/abs/2603.23495\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eAgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation\u003c/strong\u003e — uses MLLM reasoning over object tracks for zero-shot referring video object segmentation, no training required. — \u003ca href=\"https://arxiv.org/abs/2603.23489\"\u003ehttps://arxiv.org/abs/2603.23489\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ol\u003e\n\u003ch2 id=\"2-hot-github-open-source\"\u003e2. Hot GitHub Open Source\u003c/h2\u003e\n\u003ctable\u003e\n\t\u003cthead\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003cth\u003eProject\u003c/th\u003e\n\t\t\t\t\t\u003cth\u003e⭐\u003c/th\u003e\n\t\t\t\t\t\u003cth\u003eNotes\u003c/th\u003e\n\t\t\t\u003c/tr\u003e\n\t\u003c/thead\u003e\n\t\u003ctbody\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e\u003cstrong\u003ekarpathy/autoresearch\u003c/strong\u003e\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e⭐ 55.6k\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003eAI agent automated research framework — auto-runs nanochat training experiments on a single GPU; by Karpathy\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e\u003cstrong\u003egarrytan/gstack\u003c/strong\u003e\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e⭐ 46.9k\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003eGarry Tan\u0026rsquo;s Claude Code config set: 15 tool personas (CEO, design, engineering manager, QA\u0026hellip;), ready to use\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e\u003cstrong\u003epaperclipai/paperclip\u003c/strong\u003e\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e⭐ 33.0k\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003eOpen-source zero-human company orchestration framework — agent-driven fully automated business processes\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e\u003cstrong\u003eHKUDS/CLI-Anything\u003c/strong\u003e\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e⭐ 23.0k\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003eTurns any software into an agent-native CLI — universal tool interface layer, by HKU\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e\u003cstrong\u003egoogleworkspace/cli\u003c/strong\u003e\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e⭐ 22.5k\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003eOfficial Google Workspace CLI covering Drive/Gmail/Calendar/Sheets, with built-in AI agent skills\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\u003c/tbody\u003e\n\u003c/table\u003e\n\u003ch2 id=\"3-hackernews-top-posts\"\u003e3. HackerNews Top Posts\u003c/h2\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e[455pts/256c] A real time AI video agent with under 1 second of latency\u003c/strong\u003e — a real-time AI video conversation agent with \u0026lt;1s latency; a phenomenon on HN. — \u003ca href=\"https://news.ycombinator.com/item?id=41710227\"\u003ehttps://news.ycombinator.com/item?id=41710227\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003e[32pts/18c] Why are so many rolling out their own AI/LLM agent sandboxing solution?\u003c/strong\u003e — why developers build custom agent sandboxes and what a good-enough standard looks like.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ol\u003e\n\u003ch2 id=\"4-deep-reads\"\u003e4. Deep Reads\u003c/h2\u003e\n\u003ctable\u003e\n\t\u003cthead\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003cth\u003ePriority\u003c/th\u003e\n\t\t\t\t\t\u003cth\u003eItem\u003c/th\u003e\n\t\t\t\t\t\u003cth\u003eLink\u003c/th\u003e\n\t\t\t\u003c/tr\u003e\n\t\u003c/thead\u003e\n\t\u003ctbody\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e🌟\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e\u003cstrong\u003eUniGRPO\u003c/strong\u003e — unified policy optimization for visual generation reasoning\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e\u003ca href=\"https://arxiv.org/abs/2603.23500\"\u003earxiv\u003c/a\u003e\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e🌟\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e\u003cstrong\u003ekarpathy/autoresearch\u003c/strong\u003e — AI agent automated research\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e\u003ca href=\"https://github.com/karpathy/autoresearch\"\u003eGitHub\u003c/a\u003e\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\t\t\u003ctr\u003e\n\t\t\t\t\t\u003ctd\u003e💡\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e\u003cstrong\u003eFoveated Diffusion\u003c/strong\u003e — foveation-based efficient generation\u003c/td\u003e\n\t\t\t\t\t\u003ctd\u003e\u003ca href=\"https://arxiv.org/abs/2603.23491\"\u003earxiv\u003c/a\u003e\u003c/td\u003e\n\t\t\t\u003c/tr\u003e\n\t\u003c/tbody\u003e\n\u003c/table\u003e\n\u003chr\u003e\n",
  "summary": "Daily Research Brief 2026-03-26 📊 Token usage: estimated from retrieval and writing scale.\nCovers the latest AI research, open source and industry moves, updated daily.\nEditor\u0026rsquo;s Note Today\u0026rsquo;s paper thread: unified policy optimization for reasoning-driven visual generation (UniGRPO), generalized unconstrained urban 3D occupancy (OccAny), foveation-inspired efficient image/video generation, on-demand vision interaction for VLLM efficiency, and zero-shot referring video object segmentation (AgentRVOS). GitHub trending is dominated by agent infrastructure — Karpathy\u0026rsquo;s autoresearch, gstack, paperclip, CLI-Anything and Google Workspace CLI.\n"
}
