Please confirm you are human

This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.

A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.

Hold with a pointer, or hold Space or Enter.

News


thenewstack.io > ai-agent-trace-debugging

It passed CI. It passed your evals. The customer still got the wrong answer.

1+ day, 3+ hour ago   (771+ words) Your AI agent returned a 200, passed its faithfulness check, and still answered the wrong question. The evidence that explains why lives in the trace....


thenewstack.io > spec-driven-sdlc-gates

The AI-native SDLC won't be one process

2+ day, 3+ hour ago   (92+ words) Anthropic says code is no longer the bottleneck. It's right -- but the process that catches your agent's mistakes can't be one size for every change....


thenewstack.io > aws-pizza-bot-agent-inbox

AWS open-sources Pizza Bot: email-style inbox for background AI agents

3+ day, 19+ hour ago   (280+ words) Two thousand people inside Amazon used early versions - before Pizza Bot was rebuilt as a standalone community project....


thenewstack.io > stop-ai-code-sprawl

Stop AI code sprawl before it destroys your software design

4+ day, 5+ hour ago   (450+ words) Prevent AI code sprawl and Comprehension Debt. Use Python tools like pytest-archon to enforce Executable Architecture in your CI/CD....


thenewstack.io > claude-build-agents-benchmark

Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.

4+ day, 21+ hour ago   (492+ words) Sierra has open-sourced Hyper-𝜏-bench, a follow-up to its 2024 τ-bench that tests how well AI agents can build other agents....


thenewstack.io > building-trust-agentic-rag

Building trust in agentic RAG starts with evidence

1+ week, 2+ day ago   (500+ words) Agentic RAG requires clear evidence. Discover how tracking retrieval decisions, metadata, and citations builds trust in AI agent outputs....


thenewstack.io > ai-agent-evaluation-gates

AI agent evaluations are part of the product

1+ week, 3+ day ago   (886+ words) Move beyond simple AI demos. Build repeatable evaluation systems, test execution paths, and enforce strict release gates for AI agents....


thenewstack.io > ai-agents-3d-playwright

It cost $33 to build a virtual Union Square. Here's what the agents got wrong.

1+ week, 4+ day ago   (152+ words) AI coding agents built a 3D browser replica of San Francisco's Union Square in two hours, then used Playwright screenshots to catch visual errors no test could....


thenewstack.io > ai-agent-retrieval-infrastructure

Want to scale AI agents without breaking anything? Retrieval engineering is the answer.

1+ week, 4+ day ago   (23+ words) On September 24, GigaOm’s Whit Walters and Vespa.ai’s Bonnie Chase will explain what breaks when hundreds of AI agents hit retrieval at once....