AI Solutions

Top GitHub Repos for Building AI Agents in 2026

The top GitHub repos for building AI agents in 2026: agent frameworks, MCP, browser automation, memory, sandboxes, evaluation and observability tools.

GPTLabAI team 6 min read

The top GitHub repos for building AI agents in 2026 fall into five groups: agent frameworks (LangGraph, OpenAI Agents SDK, Claude Agent SDK, Pydantic AI, Google ADK, Microsoft Agent Framework, CrewAI), the Model Context Protocol SDKs and servers, browser automation (Browser Use, Playwright MCP, Stagehand), infrastructure such as memory and sandboxes, and evaluation and observability tools. You rarely need more than one or two from each group.

Every repository below was checked as of September 2026: it exists, it is actively maintained, and we note the licence where GitHub reports a standard one. We deliberately leave out star counts because they change daily and say little about fit.

How we picked these repositories

  • Active maintenance — recent commits and releases, not an abandoned demo.
  • Production use — clear docs, versioned releases, a real user base.
  • Clear licence — permissive licences are listed; others are flagged “check repo”.
  • Durability — backed by a company or a large, healthy community.

Agent frameworks

A framework handles the loop: calling the model, running tools, managing state and handing off between agents. Pick based on your language and how much control you need. Our AI agent frameworks compared post goes deeper.

Repository Language Licence Best for
langchain-ai/langgraph Python, JS MIT Explicit graph-based workflows with state, checkpoints and human-in-the-loop
openai/openai-agents-python Python MIT Lightweight multi-agent workflows with handoffs and guardrails
anthropics/claude-agent-sdk-python Python (TypeScript version also available) MIT Building agents on the same tool harness as Claude Code
pydantic/pydantic-ai Python MIT Type-safe agents with validated, structured outputs; model-agnostic
google/adk-python Python Apache 2.0 Code-first agents with built-in evaluation and deployment paths
microsoft/agent-framework Python, .NET MIT Enterprise multi-agent orchestration; the successor to AutoGen
crewAIInc/crewAI Python MIT Role-based teams of agents, quick to prototype
huggingface/smolagents Python Apache 2.0 Minimal agents that act by writing code
mastra-ai/mastra TypeScript Check repo Agents and workflows in a TypeScript/Node stack
vercel/ai TypeScript Check repo Tool calling and streaming UIs in web apps

A note on AutoGen: microsoft/autogen is now in maintenance mode, and its README directs new users to Microsoft Agent Framework. Don’t start new projects on it.

Frameworks for retrieval-heavy agents

If your agent mostly searches and reasons over documents, these are worth a look:

  • run-llama/llama_index (MIT) — document parsing, indexing and agentic retrieval.
  • deepset-ai/haystack (Apache 2.0) — pipeline-based orchestration for search and RAG.
  • stanfordnlp/dspy (MIT) — “programming, not prompting”: optimises prompts and pipelines against a metric.

See our RAG best practices for the retrieval side.

Model Context Protocol (MCP)

MCP is the open standard for connecting agents to tools and data. Write a tool server once and it works across Claude, ChatGPT, IDEs and most frameworks above.

Treat third-party MCP servers like any dependency: read the code, pin versions and limit their permissions. Our MCP servers guide covers what to install and what to avoid.

Browser automation

For agents that need to use websites without an API:

Repository Licence What it does
browser-use/browser-use MIT Python library that lets an LLM agent drive a real browser
microsoft/playwright-mcp Apache 2.0 MCP server exposing Playwright, using the accessibility tree rather than screenshots
browserbase/stagehand MIT SDK mixing natural-language actions with deterministic Playwright code

Browser agents are powerful but fragile and expensive. In our projects we prefer an official API or a deterministic Playwright script wherever one exists, and use an LLM only for the steps that genuinely change. Remember that everything on a web page is untrusted input to your agent.

Memory, state and sandboxes

  • letta-ai/letta (Apache 2.0) — a platform for stateful agents with long-term memory.
  • mem0ai/mem0 (Apache 2.0) — a memory layer you can add to existing agents and apps.
  • e2b-dev/E2B (Apache 2.0) — isolated cloud sandboxes for running agent-generated code.

Add persistent memory only when you have a clear use for it, and decide up front what the agent is allowed to remember and for how long, particularly under GDPR.

Evaluation and observability

This is the group most teams underinvest in. Without traces and evaluations, you cannot tell whether a prompt or model change made the agent better or worse.

Repository Licence Use it for
promptfoo/promptfoo MIT Test suites for prompts, agents and RAG, plus red teaming, from the CLI or CI
confident-ai/deepeval Apache 2.0 Pytest-style LLM evaluation metrics
vibrantlabsai/ragas Apache 2.0 Retrieval and RAG-specific metrics
langfuse/langfuse Check repo Self-hostable tracing, evals and prompt management
Arize-ai/phoenix Check repo Tracing and evaluation built on OpenTelemetry

Our LLM evaluation guide shows how to wire these into CI.

Coding agents you can learn from

Open-source coding agents are excellent reference implementations of tool use, planning and context management:

We compare commercial and open-source coding tools in AI coding assistants compared.

A sensible starter stack

For a typical business agent in Python, a stack we’d be comfortable putting in production:

  1. One framework — LangGraph or Pydantic AI for explicit control, or the OpenAI Agents SDK / Claude Agent SDK if you are committed to that vendor.
  2. Tools via MCP — using the official SDK, with each server’s permissions kept narrow.
  3. Tracing from day one — Langfuse or Phoenix.
  4. An evaluation suite in CI — promptfoo or DeepEval, built from real tasks.
  5. A sandbox for any code execution.

Choose the design pattern before the library. Our agentic AI design patterns guide explains which patterns need which capabilities.

Key takeaways

  • Pick one framework that matches your language and control needs; avoid stacking several.
  • MCP is the standard way to give agents tools; use the official SDKs.
  • AutoGen is in maintenance mode; new projects should look at Microsoft Agent Framework or alternatives.
  • Browser automation works, but prefer APIs and deterministic scripts where possible.
  • Evaluation and tracing tools matter as much as the framework.
  • Check licences and maintenance status before adopting, and pin versions.

From repositories to a working agent

Open-source building blocks make agents much faster to build, but choosing, securing and evaluating them still takes experience. Our AI solutions team helps companies pick a lean stack, build agents around real business processes and keep them measurable in production. If you’re planning an agent project, get in touch.

Have a project in mind? Let’s talk.

Whether you run a business or a research group, tell us what you need built, fixed or evaluated. You get a free consultation and a clear written estimate — no obligation.

  • Free consultation
  • Written scope and estimate
  • We reply within one working day
Contact us