{"version":1,"type":"rich","provider_name":"Libsyn","provider_url":"https:\/\/www.libsyn.com","height":90,"width":600,"title":"MLA 028 AI Agents: Loops, Tools, Memory, Protocols, and Evaluation","description":" What an AI agent actually is, why coding agents got good first, how memory really works, what MCP and A2A standardize, which SDKs are alive, how to evaluate on trajectories, and where the products stand after browser agents contracted. Links   Try a walking desk&amp;nbsp;- stay healthy &amp;amp; sharp while you learn &amp;amp; code  More OCDevel shows - this one has siblings, each on its own subject and produced the same way   First of two episodes on AI agents. This one is the architecture: the loop, tools and verifiable feedback, memory, the protocols (MCP, A2A, computer use), the SDK landscape, evaluation and observability, the product map, and when multiple agents help. The next episode, OpenClaw and the Personal Agent, applies it to one always-on assistant with security as the centerpiece. Coding-agent products and mechanics live in the vibe coding sequence starting at MLA 22. Agent vs workflow vs chat: the loop A chat model returns a message; a workflow is your code calling a model at fixed steps; an agent is a model that owns the control flow, choosing its next action from what it observes. That puts systems on a spectrum (chat, chat plus tools, workflows, agents) rather than in a binary, the framing Anthropic's Building Effective Agents uses. The loop itself is ReAct (Yao et al.): thought, action, observation, repeat, with the reasoning trace letting the model track and update a plan. What changed by 2026 is not the loop but the infrastructure around it, and every part of that infrastructure is an attack on per-step error compounding. Tools, function calling, and verifiable feedback Function calling: you describe tools as schemas, the model emits a structured call, your code executes it and returns the observation. The model never runs anything itself, which is the security model. Writing effective tools for agents gives the practical rules: few high-impact tools, clear namespaces, meaningful identifiers, token-efficient responses, descriptions treated as prompt engineering.  Effective context engineering for AI agents adds the overlap test: if a human cannot say which tool applies, neither can the agent. The central principle: coding agents got good first because tests and compilers give verifiable feedback that catches a bad step inside the same loop that made it. Find or manufacture the verifier before writing the prompt. Memory: context, retrieval, files, episodic &quot;Memory&quot; means four things: the context window (the only memory the model has), retrieval from an external store, files on disk, and episodic records of prior sessions. Most agent memory is files. Anthropic's  memory tool is a client-side file protocol (view, create, replace, insert, delete) against storage you own. The hard part is context management, and both labs converged on the same three mechanisms:  context editing to clear stale tool results, compaction to summarize near the limit, and notes written to files before summarization. OpenAI's Responses API conversation state has the same shape with a compaction threshold and compact endpoint. Third-party layers Mem0, Letta (from MemGPT), and Zep (temporal knowledge graph) now compete with first-party primitives. Multi-session patterns:  Effective harnesses for long-running agents. Protocols: MCP, A2A, computer use Model Context Protocol is the agent-to-tool standard, now a Linux Foundation project with individual-maintainer governance. 2026 additions: elicitation (server asks the user mid-operation), an extensions mechanism, and the async Tasks extension for long-running tools. Every major SDK below consumes it; its cost is the context each connected server's tool list occupies. A2A is the agent-to-agent standard, Google-built, Linux Foundation-hosted, at v1.0 with a steering committee spanning AWS, Cisco, Google, IBM, Microsoft, Salesforce, SAP and ServiceNow. Strong governance, weak observed consumption; worth knowing, not yet worth building on for small teams. Computer use is the universal fallback: Anthropic's  computer use tool (GA toolset with zoom and an automatic injection classifier), Google's Gemini computer use, open-source Browser Use, and Playwright MCP, which drives the accessibility tree instead of screenshots. Prefer API, then accessibility tree, then screenshots. Building one: the SDKs Both labs advise starting without a framework: Building Effective Agents and OpenAI's  A Practical Guide to Building Agents. The 2026 SDKs have converged on that critique as thin harnesses around a loop.  Claude Agent SDK: Claude Code's loop as a library (built-in tools, subagents, hooks, MCP, permissions, compaction); TypeScript and Python; pre-1.0. OpenAI Agents SDK: handoffs, guardrails, sessions, tracing on the Responses API. OpenAI deprecated the visual Agent Builder in favor of it. LangGraph and LangChain 1.x: stateful graph with checkpointing, interrupts, durable execution; create_agent as a minimal middleware harness. LangSmith is the separate tracing product. Google ADK: code-first hierarchical agent trees with native A2A; deploys to Vertex Agent Engine. Microsoft Agent Framework: GA successor to AutoGen and Semantic Kernel; Python and .NET. CrewAI: role-based crews, past 1.0, with a commercial management platform. smolagents: code agents that write Python instead of JSON calls; weakest maintenance signal on the list. Pydantic AI and Vercel AI SDK: typed validation-first agents in Python; loop control and agent abstraction in TypeScript.  Decision rule: machine-operating agent fast, Claude Agent SDK; lightweight handoffs, OpenAI; durable human-in-the-loop state, LangGraph; inside Google or Microsoft, their kit; to understand what you run, write the loop yourself first. Evaluation and observability Agents are evaluated on trajectories, not answers: traces, task evals, cost per task. Traces follow the OpenTelemetry GenAI semantic conventions; products include LangSmith, Langfuse (open source,  acquired by ClickHouse), Arize Phoenix, Braintrust, W&amp;amp;B Weave, and Helicone. Public benchmarks show the shape of a task eval: SWE-bench Verified, which  OpenAI stopped reporting citing contamination; SWE-bench Pro; tau2-bench; Terminal-Bench 2.0; OSWorld-Verified; GDPval. Cost and reliability: Princeton's Holistic Agent Leaderboard (paper) and its reliability dashboard separate pass@k capability from pass^k reliability; METR time horizons with their own limitations note. Guardrails:  OpenAI agent safety, NeMo Guardrails, Guardrails AI; prompt injection framed by Simon Willison's lethal trifecta and Google's CaMeL architectural defense. Products Claude Cowork: &quot;Claude Code for everyone,&quot; a sandboxed desktop agent with open-sourced plugins. ChatGPT agent remains; the  Atlas browser was retired within a year, folded into ChatGPT and Codex. Google discontinued Project Mariner and moved the capability into Gemini and Antigravity, which  absorbed Gemini CLI. Standalone browser agents contracted; the capability moved into models and existing apps. Still shipping: Perplexity Comet (free), Manus (ownership contested this year; check before building on it), Devin,  Copilot Studio, Agentforce. Glue: n8n (AI Agent node inside a drawn workflow, MCP server trigger) and Zapier Agents with Zapier MCP. Browser-agent prompt injection is the documented security problem: the  PleaseFix research note. Multi-agent: when it helps Two essays a day apart:  How we built our multi-agent research system (orchestrator plus parallel subagents beat a single agent on research at roughly 15x the tokens; token usage explained most of the variance) and Cognition's Don't Build Multi-Agents (dispersed decisions and unshared context make it fragile). The disagreement is task shape. Parallelize independent, read-mostly work; keep stateful, sequential work single-threaded; prefer a small hierarchy where workers return findings rather than decisions. Related episodes  MLA 29: OpenClaw and the Personal Agent MLA 22: Vibe Coding in 2026 MLA 23: Inside a Coding Agent MLA 24: Agentic Software Engineering  Companion show: Agentic Business on Gnothi follows one business as agents take on research, software, sales and operations. ","author_name":"Machine Learning Guide","author_url":"https:\/\/ocdevel.com\/mlg","html":"<iframe title=\"Libsyn Player\" style=\"border: none\" src=\"\/\/html5-player.libsyn.com\/embed\/episode\/id\/40187345\/height\/90\/theme\/custom\/thumbnail\/yes\/direction\/forward\/render-playlist\/no\/custom-color\/88AA3C\/\" height=\"90\" width=\"600\" scrolling=\"no\"  allowfullscreen webkitallowfullscreen mozallowfullscreen oallowfullscreen msallowfullscreen><\/iframe>","thumbnail_url":"https:\/\/assets.libsyn.com\/secure\/item\/40187345"}