The Agentic Stack 2026, Part 2: 20 More Tools for Memory, Quality and Execution
In short
Pinecone, Weaviate, Zep, LiteLLM, Braintrust, Composio and 14 more: the layers you only need once a prototype goes live — with limits, sources and four selection questions.

Table of Contents
Part one of this series described twenty tools that carry an agentic stack: database, delivery, runtime, observability, business processes. Afterwards the same question kept coming back: what is still missing?
Mostly what you only notice once a prototype goes into operation. Where does an agent's memory live when the context window is not enough? How do you get model cost per use case onto one report? Where does code an agent wrote itself actually run? And who says whether a change improved quality rather than just sounding different?
The following twenty tools answer exactly those questions. Sorted by category, each with its own glossary entry. We do not quote prices: they change faster than this article, and almost every vendor bills by usage.
One note up front that matters more than any tool choice: for none of these services should you assume EU data residency or data processing terms. Demonstrably self-hostable in this list are Weaviate, Qdrant, Helicone, Arize Phoenix and Trigger.dev. For everything else, settle it contractually instead of assuming it.
Contents
- Data and memory: Pinecone, Weaviate, Zep, LlamaIndex
- Model access and agent frameworks: LiteLLM, Pydantic AI, Microsoft Agent Framework
- Infrastructure and execution: Modal, Daytona, Trigger.dev
- Web data and tool access: Apify, Scrapfly, Composio
- Quality and observability: Braintrust, Helicone, Arize Phoenix
- Identity, voice and payment: Clerk, ElevenLabs, Vapi, Stripe Agentic Commerce
1) Data and memory
An agent is only as good as what it finds — and what it remembers from the last conversation.
Pinecone
Managed vector database for semantic search and RAG. Indexes, namespaces for tenant separation, metadata filters before search — without running servers. Worth it from the point where pgvector in the application database becomes noticeably slower; below that, a second database is mostly overhead.
Weaviate
Open-source vector database with hybrid search: keyword and meaning in one weighted query. That is what rescues product codes and brand names, which purely semantic search regularly loses. And because self-hosting is demonstrably possible, it stays a real option in GDPR discussions.
Zep
Memory layer that builds a temporal knowledge graph from conversations — facts as relationships with a validity period. Outdated statements are not deleted, they become invalid. The open-source core is Graphiti. Caution: this layer is deliberately a personal profile and needs purpose limitation, retention and a deletion path.
LlamaIndex
Framework for the unglamorous part of RAG: ingest, chunk, index, retrieve, cite. Most bad answers come from preparation, not the model — chunks too large, tables lost, metadata missing, no reranking. That is exactly where it works.
2) Model access and agent frameworks
The layer where model calls become a cost centre and a loop becomes a maintainable flow.
LiteLLM
100+ models in the OpenAI format, as a library or as a gateway you run. The gateway is the real reason: virtual keys per team, budgets, fallback routes on outage and — decisive for decision makers — cost per use case on one report. It is also a central chokepoint and needs redundancy.
Pydantic AI
Python agent framework with typed outputs: the model result is checked against a schema before the application uses it. For teams already using Pydantic, this is not a new way of thinking. Typing prevents format errors — not answers that are simply wrong.
Microsoft Agent Framework
The merge of AutoGen and Semantic Kernel for .NET and Python. For new projects this is the path: AutoGen is, per Microsoft, in maintenance. Convenient in Microsoft environments, and for exactly that reason a deliberate decision about lock-in.
→ Microsoft Agent Framework in the glossary
3) Infrastructure and execution
Three different questions: where does it compute, where may foreign code run, and how does a run survive a failure?
Modal
Python functions in the cloud, including on GPUs, without infrastructure work. Built for bursts: re-vectorise 200,000 passages, re-caption an image catalogue, render a video batch. Between runs nothing runs. Careful: a badly written parallel run burns a monthly budget in minutes.
Daytona
Short-lived, isolated environments for AI-generated code. As soon as an agent executes code, the answer to "where does that run?" must not be "on the application server". Important for selection: the public repository has been frozen since June 2026 — for new projects E2B is currently the more obvious choice.
Trigger.dev
Long-running background tasks in TypeScript without serverless time limits: resumable, versioned, traceable. Every serious AI feature in marketing is a background process — 240 creative variants in four languages are not an HTTP request. Open source and self-hostable.
4) Web data and tool access
Without access to the world, an agent stays a text generator.
Apify
Marketplace of ready-made scrapers (Actors), callable via API or MCP. Recurring competitor and price monitoring is bought here rather than maintaining breaking selectors yourself. The legal review of whether a source may be collected is not something a tool takes off your hands.
Scrapfly
Scraping API that handles the unpleasant part: blocking, proxies, real rendering. The effort in web access is not parsing but getting in. Technically possible does not mean permitted — robots.txt is an instruction, not an obstacle.
Composio
Tool layer giving agents access to hundreds of business applications, including per-end-user auth. That is the difference between an agent that suggests text and one that completes tasks. And for exactly that reason: writing actions behind approval steps, permissions at the minimum.
5) Quality and observability
The part that turns an experiment into a product with an acceptance criterion.
Braintrust
Evaluation platform: test datasets, scoring functions, version comparison, quality gate in delivery. Without it, every prompt change is a gut feeling. 50 to 200 real requests with a clear expected result say more than a thousand invented ones.
Helicone
Open-source observability for model calls: cost, latency, errors, usage per application, plus caching. The typical insight after four weeks is banal and expensive: a large share of cost falls on a feature used by three people. Self-hostable — relevant when prompts must not leave your network.
Arize Phoenix
Open-source tracing and evaluation on OpenTelemetry conventions. The biggest lever is retrieval diagnosis: if the right passage ranks 12th and never reaches the context, no better prompt helps — reranking does. Runs locally and self-hosted.
→ Arize Phoenix in the glossary
6) Identity, voice and payment
This is where it is decided on whose behalf an agent acts — and how far it may go.
Clerk
Sign-in, sessions, organisations and roles as a service. For agentic systems the core is not the login box but the identity question: if the agent acts with the signed-in user's permissions, damage is contained. If it acts with a service key, every access rule is bypassed.
ElevenLabs
Speech synthesis in many languages and voices, plus building blocks for voice agents. For multilingual campaigns this is the difference between one spot and twelve. Two hard conditions: voices of real people only with documented consent, and AI interaction must be identifiable.
Vapi
Voice agents on the phone: answer, listen, respond, call tools, transfer. It works for narrowly scoped tasks — appointment, callback, status. As soon as commitments, prices or contracts are involved, handover to a human is the right answer, not a failure.
Stripe Agentic Commerce
Building blocks for agents initiating and completing purchases — MCP server, approaches to agentic payments. Parts of it are explicitly preview. Our recommendation is accordingly split: make product data machine-readable now, do not build payment flows on preview features.
→ Stripe Agentic Commerce in the glossary
How to choose
Four questions, in this order, prevent most bad decisions.
Where does the data live, and who processes it? With personal data this is not a detail. Demonstrably self-hostable in this list are Weaviate, Helicone, Arize Phoenix and Trigger.dev. Everything else needs a data processing agreement covering region and logging.
What does operation cost, not the licence? Two extra services mean two more systems with backups, upgrades and on-call. A vector database alongside your existing Postgres instance needs a reason beyond "every tutorial uses one".
Is the project alive? Daytona's frozen repository and AutoGen in maintenance are the two examples in this list. Neither is a verdict on quality, but both are a reason to evaluate an alternative before a workflow depends on it.
Is quality measured? Without a test set and scoring, every improvement is a claim. This is the one category here we advise against deferring: evaluation and observability belong in the first sprint, not the fifth.
What we use ourselves
We work in production with Postgres and row level security as the data base, with preview delivery per change, with durably executed background tasks for anything longer than a request, and with tracing plus fixed test sets for every customer-facing AI feature. We only ship voice and telephony agents with clear AI disclosure and a defined handover to humans.
What we do not do: adopt tools because they appear on a list. Every additional system has to take over work someone previously did by hand — otherwise it moves work instead of saving it.
If you are currently evaluating one of these layers, let us talk about your specific case rather than tool lists. And if you have not read part one: the twenty tools of the base layer are there.
Frequently Asked Questions
What is "The Agentic Stack 2026, Part 2: 20 More Tools for Memory, Quality and Execution" about?
Pinecone, Weaviate, Zep, LiteLLM, Braintrust, Composio and 14 more: the layers you only need once a prototype goes live — with limits, sources and four selection questions.
Contents: what matters?
Data and memory: Pinecone, Weaviate, Zep, LlamaIndex Model access and agent frameworks: LiteLLM, Pydantic AI, Microsoft Agent Framework Infrastructure and execution: Modal, Daytona, Trigger.
Data and memory: what matters?
An agent is only as good as what it finds — and what it remembers from the last conversation.
Pinecone: what matters?
<img src="/blog/tools/pinecone.webp" alt="Flat illustration of a managed vector database with index cards and similarity search" loading="lazy" width="800" height="450" / Managed vector database for semantic search and RAG.
Related Articles
You might also be interested in these posts
Tools & TechnologyThe Agentic Stack 2026: 20 Tools That Carry Agentic Development
Supabase, Vercel, Inngest, Langfuse, Attio and 15 more: the five layers of an agentic stack, what each tool does, where the limits are — and four criteria for your own selection.
Tools & TechnologyLegacy Refactoring and Code Translation: What Agents Really Deliver
Python to Rust in days rather than months: how agentic porting works in practice, why factors of 10 to 40 are not transferable and which tests, benchmarks and ownership must be settled first.
Tools & TechnologyGrok Bot: xAI's AI Teammates With Their Own Computer
In beta since 11 August 2026: bots with their own cloud computer sign in to your tools and finish tasks. What that means for marketing, permissions and Grokipedia.