56 tools Avg score 8.5/10 45 with a free tier

AI agent tools

"Agent" is the least stable word in this index. Three unrelated products answer to it.

Data last refreshed 2026-07-21

Three things called the same thing

Sorting this list starts with admitting the label covers three different products. There are frontier chat models that can call tools, which are agents in the sense that they take actions on your behalf inside a conversation. There are coding agents that operate on a repository across many steps without a human in the loop for each one. And there is workflow automation — the connective software that fires model calls on triggers and moves results between systems.

A buyer for one of those is almost never a buyer for the others. Nothing about evaluating an automation platform prepares you to evaluate a repo agent, and the pricing shapes are unrelated: per seat, per run, per token, per resolved outcome.

This tag also overlaps development heavily by construction, because the most mature agents in the world right now write code. If you came here to shop for a coding agent specifically, the developer page is the better-filtered version of this list.

What carries this tag

  • Assistants that take multi-step actions rather than answering one prompt.
  • Coding agents that write, run and revise across a project.
  • Workflow and automation platforms that orchestrate model calls between apps.
  • Research agents that plan and execute their own search strategy.
  • A plain chatbot with no tool use does not carry this tag.

Which categories these come from

43 of 56 are free or freemium, 13 are paid only.

The top 12, ranked

Ranked on AI Score, then adjusted for how recently the tool shipped, whether it is an editor's pick, and what readers actually open — so a strong recent release can edge out a slightly higher score. Every entry shows what it is good for and what it costs you, including the parts the vendor leads away from.

1 Cursor logo
Cursor 9.5/10 Freemium ★ Pick Coding

AI-first code editor built on VS Code. Multi-file editing, intelligent refactoring, and a built-in chat that understands your entire codebase. A favorite among developers for complex projects.

Best for Professional developers handling complex, multi-file refactors who want AI built into a familiar VS Code-based editor.

  • Composer edits multiple files at once, handling imports and cross-file references automatically
  • Indexes the entire codebase for context-aware suggestions
  • Cmd+K inline editing and Tab completion
  • Agent mode for autonomous multi-step tasks
Free tier

2000 completions and 50 premium requests per month

Trade-offs
  • Can be resource-heavy on older machines
  • Free tier has limited premium model access
2 Claude logo
Claude 9.5/10 Free tier + Pro $20/mo + Team $30/mo/user ★ Pick Chatbots

Anthropic's AI assistant with Opus 4.6, 1M-token context, computer use, agentic coding via Claude Code, and MCP integrations.

Best for Developers and professionals who need agentic coding, computer control, and large-document or codebase analysis in one assistant.

  • 1M-token context window for entire codebases or book-length documents in one session
  • Computer use lets Claude click, type, and navigate apps directly
  • Claude Code CLI writes, tests, and iterates on code autonomously in a repo
  • MCP ecosystem connects Claude to third-party databases, APIs, and creative tools
Free tier

Free access to Claude Sonnet 4.6 with daily usage limits

Trade-offs
  • No image or video generation — strictly text and code output
  • Free tier usage limits are restrictive, especially during peak hours
3 ChatGPT logo
ChatGPT 9.5/10 Free tier + Plus $20/mo + Pro $200/mo ★ Pick Chatbots

OpenAI's flagship AI assistant with o3/o4-mini reasoning, GPT-4o, Advanced Voice, Sora video gen, Operator agent, and Deep Research — the most feature-packed chatbot available.

Best for Users who want one subscription covering reasoning, voice, vision, image and video generation, and agentic browsing in a single app.

  • Broadest feature set: voice, vision, image gen, Sora video, browsing, and agents together
  • o3 and o4-mini reasoning models for complex math, science, and coding
  • Operator agent and Deep Research add autonomous multi-step task completion
  • Free tier now includes limited GPT-4o access, not just the mini model
Free tier

Free access to GPT-4o mini and limited GPT-4o with basic features

Trade-offs
  • Pro plan at $200/mo is hard to justify unless you need heavy o3 or Operator usage
  • Writing quality has fallen behind Claude for nuanced, long-form content
4 GPT-5.5 logo
GPT-5.5 9.4/10 API: $5/$30 per 1M tokens (in/out). ChatGPT Plus $20/mo, Pro $200/mo ★ Pick Coding

OpenAI's most capable frontier model, built for complex multi-step reasoning, agentic tool use, and deep coding tasks. Powers ChatGPT and Codex with up to 272K context.

Best for Developers and teams needing a frontier reasoning model for agentic coding workflows and large-codebase context handling.

  • 272K context window ingests entire codebases or long documents without chunking workarounds
  • Agentic tool use with self-correction across multi-step workflows, not just single-turn answers
  • Powers the Codex agent for autonomous pull-request writing, testing, and iteration inside ChatGPT
  • Specialized GPT-5.5-Cyber variant tuned for cybersecurity analysis
Free tier

No free API tier. Free ChatGPT users get GPT-4o, not GPT-5.5.

Trade-offs
  • API pricing is premium — $30/M output tokens adds up fast for heavy usage
  • No free tier for the API; ChatGPT Free users are stuck on GPT-4o
5 GPT-5.6 Sol logo
GPT-5.6 Sol 9.1/10 API usage-based: Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per 1M tokens ★ Pick Chatbots

OpenAI's flagship GPT-5.6 family (Sol, Terra, Luna) targeting frontier coding, agentic tasks, and cybersecurity — currently in limited preview.

Best for Development and agentic-workflow teams evaluating a frontier model family they can route by task difficulty and token cost.

  • Three-tier family (Sol, Terra, Luna) trades capability against cost for the same generation
  • Claimed state-of-the-art results on Terminal-Bench, a real-world coding and agentic benchmark
  • Major reported cybersecurity capability gains, released via a gov-coordinated limited preview
  • Luna tier priced at $1/$6 per 1M tokens, cheap for a current-generation model
Pricing

API usage-based: Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per 1M tokens

Trade-offs
  • Launched in limited preview — not broadly available, so you can't reliably build production on it yet
  • Sol's $5/$30 pricing is premium; costs add up fast on token-heavy agentic loops
6 Gemini logo
Gemini 9.2/10 Free tier + Advanced $19.99/mo ★ Pick Chatbots

Google's multimodal AI assistant with Gemini 2.5 Pro/Flash, 2M-token context, Workspace integration, Gems agents, and native image generation.

Best for Google Workspace users who want an assistant that can read their Gmail, Drive, and Calendar while reasoning across huge documents or videos in one pass.

  • 2M-token context window, largest among major chatbots
  • Deep Google Workspace integration reads actual Gmail, Drive, and Calendar data
  • Native multimodal understanding across text, images, audio, and video
  • $19.99/mo tier bundles full 2.5 Pro access with 2TB of storage
Free tier

Gemini 2.5 Flash with image generation and basic multimodal features

Trade-offs
  • Creative writing and nuanced tone trail Claude and ChatGPT
  • Best features require Google ecosystem buy-in — less useful if you're not on Workspace
7 CodeRabbit logo
CodeRabbit 9/10 Free tier + Pro $24/user/mo, Pro Plus $48/user/mo (annual) ★ Pick Coding

AI code reviewer that posts inline, context-aware feedback on every pull request across GitHub, GitLab, Azure DevOps, and Bitbucket.

Best for Engineering teams that want automatic, context-aware review posted on every pull request across their git provider.

  • Builds full-repository context so feedback reflects how a change interacts with the rest of the codebase
  • Chains 40+ open-source linters and SAST tools into the same review pass
  • Works across GitHub, GitLab, Azure DevOps, and Bitbucket, plus IDE and CLI reviews
  • Interactive review chat lets you question, request fixes, or dismiss suggestions in the PR thread
Free tier

Permanent free tier with PR summaries and IDE/CLI reviews, plus a 14-day Pro Plus trial that needs no card.

Trade-offs
  • Per-PR-author billing at $24–$48/user/mo adds up fast for larger engineering teams
  • AI review comments can still be noisy or surface false positives that reviewers must triage
8 GLM-5.2 logo
GLM-5.2 8.7/10 Free 20M tokens on signup + open weights; paid API & coding plans (check site) ★ Pick Coding

Open-weight 744B-parameter frontier model from Zhipu AI built for coding, reasoning, and long-horizon agentic work, with a 1M-token context and an MIT license.

Best for Development teams building coding agents who want a frontier-class model they can self-host, fine-tune, and deploy under a fully permissive license.

  • 744B-parameter model released as MIT-licensed open weights
  • 1M-token context window for whole-codebase and long agent-history reasoning
  • Tuned for long-horizon, multi-turn agentic coding rather than one-shot completions
  • Free to start via chat UI, 20M signup tokens, or self-hosted weights
Free tier

Free chat access plus 20M free tokens on API signup, and open weights available to download and self-host under MIT.

Trade-offs
  • At 744B parameters, self-hosting demands serious GPU hardware — out of reach for most individual developers
  • Vendor-published benchmarks need independent verification; real-world coding performance may vary from the claims
9 Retell AI logo
Retell AI 8.9/10 Free $10 credit, then ~$0.07/min usage-based ★ Pick Chatbots

A developer platform for building production-grade voice AI agents that handle phone conversations with human-like latency and natural turn-taking.

Best for Developer teams building production phone-based voice AI agents for customer service at scale.

  • Sub-800ms latency for natural, human-like turn-taking
  • HIPAA-compliant, built for regulated production call volumes
  • Bring-your-own-LLM across OpenAI, Anthropic, or custom models
  • Native telephony integration with Twilio, Vonage, and SIP trunks
Free tier

$10 free credit to test the platform — enough for roughly 140 minutes of calls

Trade-offs
  • Requires developer skills — no true no-code builder for non-technical users
  • Usage-based pricing can get expensive at high call volumes without enterprise negotiation
10 Kimi K3 logo
Kimi K3 8.5/10 Free tier + paid from $19/mo; API $3/$15 per M tokens ★ Pick Chatbots

Moonshot AI's 2.8T-parameter multimodal chatbot with a 1M-token context window, tuned for coding and long-horizon agentic tasks.

Best for Developers and power users who want a long-context, coding-focused chatbot with lower token costs than comparable frontier tiers.

  • 1M-token context window for whole-codebase and long-document reasoning
  • API priced at $3 in / $15 out per M tokens, undercutting comparable US frontier tiers
  • Multimodal input paired with an OpenAI-compatible API for existing tooling
  • Topped several frontend-coding leaderboards at its July 2026 launch
Free tier

Free access to Kimi K3 with limits on messages and long-context usage.

Trade-offs
  • Launch-day benchmark leadership rarely holds as rivals ship and independent testing catches up
  • Chinese-lab data handling, content policies, and regional availability may be dealbreakers for some teams
11 Bonsai 27B logo
Bonsai 27B 8.5/10 Free — open-source weights under Apache 2.0 ★ Pick Research

Open-source 27B multimodal model with 1-bit and ternary variants that run locally on a phone or laptop for on-device agentic workflows.

Best for Developers and researchers building offline or privacy-sensitive agentic apps who need a multimodal model small enough to run locally on a phone or laptop.

  • Ternary and 1-bit weight variants, an unusual first-class release for a 27B-class model
  • Runs natively on-device on phones and laptops instead of requiring server GPUs
  • Apache 2.0 license permits commercial use, fine-tuning, and redistribution with no restrictions
  • Multimodal rather than text-only at a size class previously reserved for the cloud
Free tier

Entirely free and open-source — weights released under Apache 2.0 with no paid tier from PrismML.

Trade-offs
  • 'Near full-precision' claims at 1-bit/ternary are the vendor's own and need independent benchmarking before you trust them
  • Running a 27B model on a phone still taxes RAM, thermals, and battery — real-world throughput on older devices is unproven
12 HeyGen logo
HeyGen 8.8/10 Free tier + Creator $29/mo ★ Pick Video

AI video platform that turns text scripts into realistic avatar videos with lip-sync, multilingual translation, and personalized video at scale.

Best for Marketing and sales teams producing talking-head avatar videos, multilingual dubs, or personalized outreach clips at scale.

  • Video Translate clones the speaker's voice and re-syncs lips across 40+ languages
  • Personalized Video API generates thousands of individualized outreach videos from one template
  • Interactive Avatar supports real-time conversational use cases like AI receptionists
Free tier

Free tier lets you test avatar creation with 1 credit — enough to evaluate quality but too limited for real production

Trade-offs
  • Free tier is extremely limited — essentially a demo with 1 credit
  • Stock avatar quality trails Synthesia's top-tier options slightly

The other 44 in the index

Same tag, lower down the ranking. Scores, pricing and full write-ups behind each name.

Vapi Chatbots · Free tier, then ~$0.05/min + provider costs 8.8/10 GitHub Copilot Coding · Free tier + Pro $10/mo 8.8/10 Windsurf Coding · Freemium 9.1/10 Fathom Productivity · Freemium — Free core; Team $19/user/mo 8.8/10 Zed Coding · Free (open source) — BYOK for AI features 8.8/10 Grok Chatbots · Free tier + SuperGrok $30/mo + Heavy $300/mo 8.7/10 Writer Writing · From $18/user/mo, Enterprise custom 8.7/10 Fireflies.ai Productivity · Freemium — Pro $10/user/mo, Business $19/user/mo 8.7/10 Bland AI Chatbots · From $0.09/min or plans from ~$499/mo 8.6/10 Sierra Chatbots · Enterprise only — outcome-based, contact sales 8.7/10 Zapier Productivity · Freemium 8.9/10 Replit Agent Coding · Freemium 8.8/10 WorkBeaver Productivity · Free (25 tasks one-time) + Priority $23.95/mo + Founder $108/mo 8.5/10 Otter.ai Productivity · Freemium — Free (300 min/mo), Pro $8.33/user/mo, Business $20/user/mo 8.5/10 Glean Productivity · Custom enterprise pricing 8.8/10 Laguna XS.2 Coding · Free (Apache 2.0) 8.5/10 n8n Productivity · Freemium 8.7/10 OpenClaw Productivity · Free (Open Source) 8.6/10 MiniMax M3 Coding · Freemium — Plus $20/mo, API $0.30/M input / $1.20/M output 8.5/10 Fugu-Cyber Research · Token plan only: $6–$12/M input, $36–$54/M output (application required, no free tier) 8.2/10 Devin Coding · Teams ~$500/user/mo, Enterprise custom 8.5/10 Command A+ Coding · Free API tier + open weights (Apache 2.0) 8.5/10 Muse Spark 1.1 Coding · Free in Meta AI app; API in public preview (pricing not yet disclosed) 8.2/10 Parallel Search Turbo Research · API usage-based, Turbo from $1 per 1,000 requests 8.2/10 xAI Voice Agent Builder Chatbots · Usage-based from $0.05/min 8.2/10 OpenScience Research · Free — open-source (self-hosted; you pay your own model API costs) 8.2/10 Hume AI Chatbots · Free credits + usage-based from $0.07/min 8.4/10 Siri AI Chatbots · Free with iOS 27 / macOS 27 8.3/10 Fundraisly Productivity · Custom plans (contact for pricing) 8.2/10 North Mini Code Coding · Free (open-source, Apache 2.0) 8.2/10 ZONOS2 Music · Free (open-source) + paid cloud tiers 8.2/10 Goose Coding · Free (open-source, Apache 2.0) 8.2/10 Copy.ai Writing · Freemium 8.3/10 Co-Scientist Research · Free (experimental access via registration) 8.2/10 T3MP3ST Coding · Free, open source (AGPL-3.0) — you pay for the underlying AI agent's API usage 7.8/10 Jobbie Productivity · Free to start (no credit card) + paid plans for higher volume 7.8/10 Cofounder 2 Productivity · Free (open beta — pricing TBA) 8/10 You.com Research · Free tier + YouPro $20/mo 8/10 OpenYabby Productivity · Free — open-source, self-hosted (bring your own model API keys) 7.7/10 Osaurus Coding · Free (open-source) 7.5/10 SellerClaw Productivity · Freemium (public beta) 7.8/10 Grok Build Coding · Included with SuperGrok ($30/mo) or X Premium+ ($40/mo) 7.8/10 Overtone Chatbots · Freemium — public tiers not yet announced 7.3/10 Microsoft Scout Productivity · Private preview (Frontier program + GitHub Copilot license required) 7.5/10

Open all 56 in the filterable directory →

Where to start without paying

45 of the 56 tools here publish a free tier. These are the strongest of them, with what the free tier actually gets you.

Cursor Free tier

2000 completions and 50 premium requests per month

Claude Free tier

Free access to Claude Sonnet 4.6 with daily usage limits

ChatGPT Free tier

Free access to GPT-4o mini and limited GPT-4o with basic features

GPT-5.5 Free tier

No free API tier. Free ChatGPT users get GPT-4o, not GPT-5.5.

Gemini Free tier

Gemini 2.5 Flash with image generation and basic multimodal features

CodeRabbit Free tier

Permanent free tier with PR summaries and IDE/CLI reviews, plus a 14-day Pro Plus trial that needs no card.

See all 45 free and freemium options →

The questions autonomy raises

What can it actually change

Read-only agents are a different risk class to agents with write access to a repo, an inbox, a CRM or a payment system. Establish the blast radius before the capability list.

Where credentials live

Agents need keys to be useful. Whether those keys sit in a vendor database, in your own infrastructure, or scoped per action is a security decision, not a feature preference.

Is there a stop and an undo

Long chains fail in the middle. Look for step-level visibility, an interrupt, and a way to roll back partial work. Plenty of tools have none of the three.

Cost per run at real volume

Multi-step agents multiply token spend by the number of steps and the number of retries. A cheap-looking per-call price becomes a different number once a run is forty calls.

Where autonomy actually breaks

  • Reliability compounds in the wrong direction. A step that succeeds 95% of the time succeeds about 60% of the time across ten steps, and most real workflows are longer than ten steps.
  • Demos are single-happy-path. The interesting question is what happens on the branch the vendor did not record, and the answer is usually that the agent confidently does the wrong thing.
  • Silent partial completion again — the most expensive failure mode in the category, because it looks exactly like success until somebody checks.
Comparison explorer Put the top three side by side Opens with Cursor, Claude, ChatGPT already loaded. Swap any of them out and read pricing plans, features, pros and cons in one table.

Related reading

Browse another job