AI Automation Research & Market Insights
Data-informed perspectives on the AI agent market, automation trends, and what they mean for how non-technical professionals and lean teams get work done.
- 31 Research LoopX: A Control Plane for Agent Loops
LoopX is a local state kernel that keeps objectives, gates, evidence and quota stable while Codex or Claude Code runs. What it does, and what it refuses to do.
- 32 Research AI Agent Evaluation: The 56% Pass Ceiling
AI agent evaluation on 107 real business tasks: the best model passes 56.1%, and the same model shifts 7 points depending on which harness runs it.
- 33 Research What Is reverse-skill? A Security Skill Router
reverse-skill hit 14K GitHub stars routing AI agents through security work. What it is, why the 10K-in-a-day claim is wrong, and the authorization line.
- 34 Research What Is Numbat? Perplexity Agent Guard
Numbat is Perplexity's open-source agent security suite for coding agents. How hooks, 52 rules, and monitor-only defaults work after the HF incident.
- 35 Research OptMem: Permanent Memory in 426 Tokens
OptMem stores agent memory in plain files instead of a vector store, in a 426-token prompt. What that buys you, what it costs, and where it breaks.
- 36 Research Alibaba's open-code-review, Explained
Alibaba open-code-review is free and Apache-2.0. What the hybrid rules-plus-LLM design buys you, what it actually costs to run, and who it fits.
- 37 Research Kimi K3 Technical Report: What It Reveals
Moonshot's 47-page Kimi K3 technical report: the KDA architecture, a WebDev Arena first, real cost curves, and a cyber eval most coverage skipped.
- 38 Research Humans and AI Agents, Working in Parallel
In one week, OpenWorker, Buzz, and ego-lite all shipped the same idea: humans and AI agents working in parallel. What the pattern means and where it goes next.
- 39 Research Opus 5 Code Review: Precise but Noisier
In CodeRabbit's own test, Opus 5 wrote more precise review comments but caught fewer known bugs and 4x the nitpicks. What that means for your code review agent.
- 40 Research Opus 5's ARC-AGI-3 Jump Didn't Transfer
Opus 5's ~4x ARC-AGI-3 lead collapses to a statistical tie with Kimi K3 and Fable 5 on a held-out suite. What the Witness benchmark says about picking a model.