Skip to main contentSkip to navigationSkip to footer
    Tools & Technology

    The Agentic Stack 2026, Part 2: 20 More Tools for Memory, Quality and Execution

    In short

    Pinecone, Weaviate, Zep, LiteLLM, Braintrust, Composio and 14 more: the layers you only need once a prototype goes live — with limits, sources and four selection questions.

    September 20, 20268 min readNick Meyer
    Share:
    The Agentic Stack 2026, Part 2: 20 More Tools for Memory, Quality and Execution

    Table of Contents

    Part one of this series described twenty tools that carry an agentic stack: database, delivery, runtime, observability, business processes. Afterwards the same question kept coming back: what is still missing?

    Mostly what you only notice once a prototype goes into operation. Where does an agent's memory live when the context window is not enough? How do you get model cost per use case onto one report? Where does code an agent wrote itself actually run? And who says whether a change improved quality rather than just sounding different?

    The following twenty tools answer exactly those questions. Sorted by category, each with its own glossary entry. We do not quote prices: they change faster than this article, and almost every vendor bills by usage.

    One note up front that matters more than any tool choice: for none of these services should you assume EU data residency or data processing terms. Demonstrably self-hostable in this list are Weaviate, Qdrant, Helicone, Arize Phoenix and Trigger.dev. For everything else, settle it contractually instead of assuming it.


    Contents

    1. Data and memory: Pinecone, Weaviate, Zep, LlamaIndex
    2. Model access and agent frameworks: LiteLLM, Pydantic AI, Microsoft Agent Framework
    3. Infrastructure and execution: Modal, Daytona, Trigger.dev
    4. Web data and tool access: Apify, Scrapfly, Composio
    5. Quality and observability: Braintrust, Helicone, Arize Phoenix
    6. Identity, voice and payment: Clerk, ElevenLabs, Vapi, Stripe Agentic Commerce

    1) Data and memory

    An agent is only as good as what it finds — and what it remembers from the last conversation.

    Pinecone

    Flat illustration of a managed vector database with index cards and similarity search

    Managed vector database for semantic search and RAG. Indexes, namespaces for tenant separation, metadata filters before search — without running servers. Worth it from the point where pgvector in the application database becomes noticeably slower; below that, a second database is mostly overhead.

    → Pinecone in the glossary

    Weaviate

    Flat illustration of hybrid search combining keyword and semantic hits

    Open-source vector database with hybrid search: keyword and meaning in one weighted query. That is what rescues product codes and brand names, which purely semantic search regularly loses. And because self-hosting is demonstrably possible, it stays a real option in GDPR discussions.

    → Weaviate in the glossary

    Zep

    Flat illustration of a temporal knowledge graph with nodes and validity periods

    Memory layer that builds a temporal knowledge graph from conversations — facts as relationships with a validity period. Outdated statements are not deleted, they become invalid. The open-source core is Graphiti. Caution: this layer is deliberately a personal profile and needs purpose limitation, retention and a deletion path.

    → Zep in the glossary

    LlamaIndex

    Flat illustration of a document pipeline from PDF to indexed chunks

    Framework for the unglamorous part of RAG: ingest, chunk, index, retrieve, cite. Most bad answers come from preparation, not the model — chunks too large, tables lost, metadata missing, no reranking. That is exactly where it works.

    → LlamaIndex in the glossary


    2) Model access and agent frameworks

    The layer where model calls become a cost centre and a loop becomes a maintainable flow.

    LiteLLM

    Flat illustration of a gateway routing requests to several model providers

    100+ models in the OpenAI format, as a library or as a gateway you run. The gateway is the real reason: virtual keys per team, budgets, fallback routes on outage and — decisive for decision makers — cost per use case on one report. It is also a central chokepoint and needs redundancy.

    → LiteLLM in the glossary

    Pydantic AI

    Flat illustration of a validated data object with a checkmark and schema fields

    Python agent framework with typed outputs: the model result is checked against a schema before the application uses it. For teams already using Pydantic, this is not a new way of thinking. Typing prevents format errors — not answers that are simply wrong.

    → Pydantic AI in the glossary

    Microsoft Agent Framework

    Flat illustration of several cooperating agents with an approval step

    The merge of AutoGen and Semantic Kernel for .NET and Python. For new projects this is the path: AutoGen is, per Microsoft, in maintenance. Convenient in Microsoft environments, and for exactly that reason a deliberate decision about lock-in.

    → Microsoft Agent Framework in the glossary


    3) Infrastructure and execution

    Three different questions: where does it compute, where may foreign code run, and how does a run survive a failure?

    Modal

    Flat illustration of serverless GPU execution with parallel instances

    Python functions in the cloud, including on GPUs, without infrastructure work. Built for bursts: re-vectorise 200,000 passages, re-caption an image catalogue, render a video batch. Between runs nothing runs. Careful: a badly written parallel run burns a monthly budget in minutes.

    → Modal in the glossary

    Daytona

    Flat illustration of an isolated sandbox environment with a shield

    Short-lived, isolated environments for AI-generated code. As soon as an agent executes code, the answer to "where does that run?" must not be "on the application server". Important for selection: the public repository has been frozen since June 2026 — for new projects E2B is currently the more obvious choice.

    → Daytona in the glossary

    Trigger.dev

    Flat illustration of a long-running task with steps and retries

    Long-running background tasks in TypeScript without serverless time limits: resumable, versioned, traceable. Every serious AI feature in marketing is a background process — 240 creative variants in four languages are not an HTTP request. Open source and self-hostable.

    → Trigger.dev in the glossary


    4) Web data and tool access

    Without access to the world, an agent stays a text generator.

    Apify

    Flat illustration of a marketplace of ready-made scraping programs

    Marketplace of ready-made scrapers (Actors), callable via API or MCP. Recurring competitor and price monitoring is bought here rather than maintaining breaking selectors yourself. The legal review of whether a source may be collected is not something a tool takes off your hands.

    → Apify in the glossary

    Scrapfly

    Flat illustration of a cloud browser with proxy rotation

    Scraping API that handles the unpleasant part: blocking, proxies, real rendering. The effort in web access is not parsing but getting in. Technically possible does not mean permitted — robots.txt is an instruction, not an obstacle.

    → Scrapfly in the glossary

    Composio

    Flat illustration of a tool layer connecting to business applications

    Tool layer giving agents access to hundreds of business applications, including per-end-user auth. That is the difference between an agent that suggests text and one that completes tasks. And for exactly that reason: writing actions behind approval steps, permissions at the minimum.

    → Composio in the glossary


    5) Quality and observability

    The part that turns an experiment into a product with an acceptance criterion.

    Braintrust

    Flat illustration comparing two variants on a scoring scale

    Evaluation platform: test datasets, scoring functions, version comparison, quality gate in delivery. Without it, every prompt change is a gut feeling. 50 to 200 real requests with a clear expected result say more than a thousand invented ones.

    → Braintrust in the glossary

    Helicone

    Flat illustration of a cost dashboard for model calls

    Open-source observability for model calls: cost, latency, errors, usage per application, plus caching. The typical insight after four weeks is banal and expensive: a large share of cost falls on a feature used by three people. Self-hostable — relevant when prompts must not leave your network.

    → Helicone in the glossary

    Arize Phoenix

    Flat illustration of a trace chain with a faulty retrieval step

    Open-source tracing and evaluation on OpenTelemetry conventions. The biggest lever is retrieval diagnosis: if the right passage ranks 12th and never reaches the context, no better prompt helps — reranking does. Runs locally and self-hosted.

    → Arize Phoenix in the glossary


    6) Identity, voice and payment

    This is where it is decided on whose behalf an agent acts — and how far it may go.

    Clerk

    Flat illustration of sign-in with organisations and roles

    Sign-in, sessions, organisations and roles as a service. For agentic systems the core is not the login box but the identity question: if the agent acts with the signed-in user's permissions, damage is contained. If it acts with a service key, every access rule is bypassed.

    → Clerk in the glossary

    ElevenLabs

    Flat illustration of a speech waveform with several language variants

    Speech synthesis in many languages and voices, plus building blocks for voice agents. For multilingual campaigns this is the difference between one spot and twelve. Two hard conditions: voices of real people only with documented consent, and AI interaction must be identifiable.

    → ElevenLabs in the glossary

    Vapi

    Flat illustration of a telephony agent handing over to a human

    Voice agents on the phone: answer, listen, respond, call tools, transfer. It works for narrowly scoped tasks — appointment, callback, status. As soon as commitments, prices or contracts are involved, handover to a human is the right answer, not a failure.

    → Vapi in the glossary

    Stripe Agentic Commerce

    Flat illustration of an agent with a basket and an approval step

    Building blocks for agents initiating and completing purchases — MCP server, approaches to agentic payments. Parts of it are explicitly preview. Our recommendation is accordingly split: make product data machine-readable now, do not build payment flows on preview features.

    → Stripe Agentic Commerce in the glossary


    How to choose

    Four questions, in this order, prevent most bad decisions.

    Where does the data live, and who processes it? With personal data this is not a detail. Demonstrably self-hostable in this list are Weaviate, Helicone, Arize Phoenix and Trigger.dev. Everything else needs a data processing agreement covering region and logging.

    What does operation cost, not the licence? Two extra services mean two more systems with backups, upgrades and on-call. A vector database alongside your existing Postgres instance needs a reason beyond "every tutorial uses one".

    Is the project alive? Daytona's frozen repository and AutoGen in maintenance are the two examples in this list. Neither is a verdict on quality, but both are a reason to evaluate an alternative before a workflow depends on it.

    Is quality measured? Without a test set and scoring, every improvement is a claim. This is the one category here we advise against deferring: evaluation and observability belong in the first sprint, not the fifth.


    What we use ourselves

    We work in production with Postgres and row level security as the data base, with preview delivery per change, with durably executed background tasks for anything longer than a request, and with tracing plus fixed test sets for every customer-facing AI feature. We only ship voice and telephony agents with clear AI disclosure and a defined handover to humans.

    What we do not do: adopt tools because they appear on a list. Every additional system has to take over work someone previously did by hand — otherwise it moves work instead of saving it.

    If you are currently evaluating one of these layers, let us talk about your specific case rather than tool lists. And if you have not read part one: the twenty tools of the base layer are there.

    Frequently Asked Questions

    What is "The Agentic Stack 2026, Part 2: 20 More Tools for Memory, Quality and Execution" about?

    Pinecone, Weaviate, Zep, LiteLLM, Braintrust, Composio and 14 more: the layers you only need once a prototype goes live — with limits, sources and four selection questions.

    Contents: what matters?

    Data and memory: Pinecone, Weaviate, Zep, LlamaIndex Model access and agent frameworks: LiteLLM, Pydantic AI, Microsoft Agent Framework Infrastructure and execution: Modal, Daytona, Trigger.

    Data and memory: what matters?

    An agent is only as good as what it finds — and what it remembers from the last conversation.

    Pinecone: what matters?

    <img src="/blog/tools/pinecone.webp" alt="Flat illustration of a managed vector database with index cards and similarity search" loading="lazy" width="800" height="450" / Managed vector database for semantic search and RAG.