Insights
Founder and builder perspectives on AI tools, thinking patterns, and the new way of working
Showing 97-108 of 485
Picking MCP Servers You Can Actually Trust in 2026
With 82% of audited MCP servers vulnerable to path traversal, picking one to connect to your systems needs a real checklist. Here's a practical one.
Why Vendor Neutrality Mattered More Than Anyone Expected
MCP's Linux Foundation handoff wasn't symbolic. It changed how competing labs, enterprises, and long-tail builders actually behaved. Here's the mechanism.
Business Application Servers: MCP's Enterprise Second Wave
950+ of MCP's 9,652 registry servers are built for customer service, sales, and internal ops. That's the signal a protocol is maturing past hobbyist adoption.
The MCP Security Audit Nobody Wanted to Read (82% Path Traversal)
A 2026 audit of 2,600+ MCP servers found 82% vulnerable to path traversal, 67% to code injection. Why it happened, and what a safe server looks like.
9,652 Servers and Counting: What's Actually in the MCP Registry
The MCP Registry hit 9,652 latest server records by May 2026. A tour of what's actually in there — the enterprise wave, the long tail, and what it means.
Inside the Stateless Rewrite: Why MCP Had to Give Up Persistent Connections
The 2026-07-28 MCP spec traded persistent connections for a stateless architecture. Here's the scaling problem that forced the trade, and what it cost.
What 97 Million Downloads a Month Actually Tells You
MCP's SDKs pull ~97M monthly downloads. That number is real, impressive, and mostly the wrong thing to be looking at. Here's what to check instead.
MCP Went From Anthropic Side Project to Linux Foundation Standard
In 14 months MCP went from an internal Anthropic spec to a Linux Foundation standard. Here's why handing away control was the move that made it win.
An Eval Suite Starter Kit for Your First Production Agent
A concrete, ordered starting eval stack for a small team shipping their first production agent — what to build first, and what to skip until it hurts.
What Correlating Judge Scores to Human Ratings Actually Looks Like
The 0.85 correlation threshold everyone cites is easy to state and tedious to actually reach. Here's the real, unglamorous process of validating a judge.
Turning Evals Into Guardrails That Run at Inference Time
Offline eval suites catch what already shipped. Runtime guardrails catch bad outputs before a user ever sees them. Here's how the two connect.
Why Eval-Driven Development Is Replacing Vibe-Checking Outputs
Reading five outputs and deciding a prompt 'feels better' doesn't scale past the first week. Eval-driven development writes the test before the fix.