Tag: ai-agents
All the articles with the tag "ai-agents".
What 639,000 Execution Steps Taught Me About How AI Agents Really Fail
Published: at 06:00 PMI applied the MAST failure taxonomy to 639,000 execution steps from AI agents running in production for five months. My first headline finding turned out to be an infrastructure bug masquerading as agent behavior. This post is about what agents actually fail at in production, and the discipline it takes to not fool yourself with production data.
O Que 639.000 Passos de Execução Me Ensinaram Sobre Como Agentes de IA Realmente Falham
Published: at 06:00 PMApliquei a taxonomia de falhas MAST a 639.000 passos de execução de agentes de IA rodando em produção por cinco meses. Meu primeiro headline acabou sendo um bug de infraestrutura disfarçado de comportamento do agente. Este post é sobre o que os agentes de fato falham em produção, e sobre a disciplina necessária para não enganar a si mesmo com dados de produção.
LHC v0.2: A Benchmark for Long-Horizon Agent Coherence (and the Methodology That Got It Honest)
Published: at 08:00 PMI just published LHC v0.2, an open benchmark for long-horizon coherence in 8B-class agent models, plus a deterministic parser baseline that puts a useful floor on what fine-tuning is worth for structured-state tasks. This post explains what they're for, how to use them, and the methodology arc that produced them across five rounds of external review.
LHC v0.2: Um Benchmark para Coerência de Longo Horizonte em Agentes (e a Metodologia que Tornou os Resultados Honestos)
Published: at 08:00 PMAcabei de publicar o LHC v0.2, um benchmark aberto para coerência de longo horizonte em modelos de agentes da classe 8B, mais um baseline de parser determinístico que coloca um piso útil sobre o que fine-tuning vale para tarefas de estado estruturado. Este post explica para que servem, como usá-los, e o arco metodológico que os produziu ao longo de cinco rodadas de revisão externa.
We're Mistaking the Bootstrap Phase for the Future of AI Agents
Published: at 08:00 AMThe self-hosted AI agent movement is real and important. But we are confusing a bootstrap phase with a destination architecture. The long-term future of agents will be defined by platforms that make them reliable, governable, and operationally boring.