AI Automation Research & Market Insights

Data-informed perspectives on the AI agent market, automation trends, and what they mean for how non-technical professionals and lean teams get work done.

  1. 31
    Research LoopX: A Control Plane for Agent Loops

    LoopX is a local state kernel that keeps objectives, gates, evidence and quota stable while Codex or Claude Code runs. What it does, and what it refuses to do.

    August 6, 2026 · 9 min read
  2. 32
    Research AI Agent Evaluation: The 56% Pass Ceiling

    AI agent evaluation on 107 real business tasks: the best model passes 56.1%, and the same model shifts 7 points depending on which harness runs it.

    August 5, 2026 · 10 min read
  3. 33
    Research What Is reverse-skill? A Security Skill Router

    reverse-skill hit 14K GitHub stars routing AI agents through security work. What it is, why the 10K-in-a-day claim is wrong, and the authorization line.

    August 3, 2026 · 13 min read
  4. 34
    Research What Is Numbat? Perplexity Agent Guard

    Numbat is Perplexity's open-source agent security suite for coding agents. How hooks, 52 rules, and monitor-only defaults work after the HF incident.

    July 31, 2026 · 9 min read
  5. 35
    Research OptMem: Permanent Memory in 426 Tokens

    OptMem stores agent memory in plain files instead of a vector store, in a 426-token prompt. What that buys you, what it costs, and where it breaks.

    July 30, 2026 · 8 min read
  6. 36
    Research Alibaba's open-code-review, Explained

    Alibaba open-code-review is free and Apache-2.0. What the hybrid rules-plus-LLM design buys you, what it actually costs to run, and who it fits.

    July 30, 2026 · 8 min read
  7. 37
    Research Kimi K3 Technical Report: What It Reveals

    Moonshot's 47-page Kimi K3 technical report: the KDA architecture, a WebDev Arena first, real cost curves, and a cyber eval most coverage skipped.

    July 29, 2026 · 11 min read
  8. 38
    Research Humans and AI Agents, Working in Parallel

    In one week, OpenWorker, Buzz, and ego-lite all shipped the same idea: humans and AI agents working in parallel. What the pattern means and where it goes next.

    July 28, 2026 · 11 min read
  9. 39
    Research Opus 5 Code Review: Precise but Noisier

    In CodeRabbit's own test, Opus 5 wrote more precise review comments but caught fewer known bugs and 4x the nitpicks. What that means for your code review agent.

    July 28, 2026 · 8 min read
  10. 40
    Research Opus 5's ARC-AGI-3 Jump Didn't Transfer

    Opus 5's ~4x ARC-AGI-3 lead collapses to a statistical tie with Kimi K3 and Fable 5 on a held-out suite. What the Witness benchmark says about picking a model.

    July 28, 2026 · 7 min read