Framework · Dashboard · Platform

Build AI in Rails.
See everything it does.

Active Agent is the framework: agents are controllers, prompts are views, tools are methods. Action Agent is the dashboard that mounts beside it, so every generation lands as a trace in your own app. Point the same telemetry at activeagents.ai when the team needs it in production.

$ bundle add activeagent actionagent · $ rails g action_agent:install && rails db:migrate
  • *MIT licensed
  • #Rails 7.2 · 8.0 · 8.1
  • @Ruby 3.2+
  • {}10 providers
  • ->Telemetry built in
@ app/agents/support_agent.rb ruby
class SupportAgent < ApplicationAgent
  generate_with :anthropic, model: "claude-sonnet-4-5"

  # current_user is whoever made the call
  before_action :authorize_tickets!
  delegate_to RefundAgent, budget: { max_calls: 1 }

  def reply
    prompt(
      message: params[:message],
      tools: TicketTools.tool_definitions,
      mcps: [ { url: "https://mcp.stripe.com" } ]
    )
  end

  def find_tickets(**filters)
    TicketTools.call("find_tickets",
                     actor: current_user, **filters)
  end
end

# system prompt: app/views/agents/support/instructions.md
SupportAgent.as(current_user)
            .with(message: "Where is my order 4821?")
            .reply.generate_now
-> Traces · SupportAgent#reply live
> -> tr_9f2e41 SupportAgent#reply 2.4s~$0.0182 200
0ms600ms1.2s1.8s2.4s
SupportAgent#reply 2.4s
├prompt 38ms
├anthropic.messages 1.1s
└thinking 340ms
├find_tickets 212ms
├anthropic.messages 0.9s
└response 21ms
T ↓ 4,812 ↑ 388 · 1,240 cached · 612 thinking @ user/4821 v7 3f9c1a finish stop
Context at call41.7k / 200k · 21%
messages 18.2k tool results 9.4k instructions 1.8k tool schemas 2.1k MCP schemas 10.2k
How it fits together

Three gems. One wire format.

The framework reports what every agent did. The dashboard stores and shows it. The platform is that dashboard, hosted. The same YAML block moves a trace from one to the next.

@ activeagent
The framework · MIT

Agents, actions and prompt templates. Ten providers, tools, MCP servers, delegation, structured output and streaming. Evaluations and telemetry ship in the gem.

gem "activeagent" rails g active_agent:install
local_storage: true
-> actionagent
The dashboard · MIT

A mountable Rails engine. Traces with span waterfalls, metrics, interactions, evaluations, the agent builder, the Run Agent workbench, and your agents served as an MCP server. On your database.

gem "actionagent" mount ActionAgent::Engine
endpoint: api.activeagents.ai
# activeagents.ai
The platform · Hosted

The same engine, multi-tenant. Workspaces and seats, retention and quotas by plan, published evaluation reports, managed sandboxes. Every workspace starts with a free trial.

api_key: workspace key POST /v1/traces

Also in the family: solid_agent persists contexts, memory, runs and cost as models you own. activeagents-telemetry reports RubyLLM apps on the same wire format.

The framework · activeagent

Agents are controllers.

Actions, callbacks, params and views, for AI. Tools are Ruby methods, MCP servers are a declaration, and another agent is just another tool. Everything below is in the MIT gem.

@Agents are controllers

Actions, before_action, params. Run now, or later on Active Job.

class TranslationAgent < ApplicationAgent
  generate_with :openai, model: "gpt-4o-mini"
  before_action :load_glossary

  def translate
    prompt   # renders translate.md.erb
  end
end

TranslationAgent.with(text: "Hello", locale: "ja")
                .translate.generate_now
TranslationAgent.with(text: "Hello", locale: "ja")
                .translate.generate_later   # Active Job
docs/agents ->

*Action Prompt

Prompts are views. ERB, partials and layouts, plus an instructions file per agent.

app/views/agents/translation/
  instructions.md      # the system prompt
  translate.md.erb     # the user turn
  translate.json       # a response schema, if any

<%# translate.md.erb %>
Translate into <%= params[:locale] %>, keeping the tone:

<%= params[:text] %>
docs/actions ->

>>Any provider

OpenAI, Anthropic, Gemini, Bedrock, Azure, Ollama, OpenRouter, Requesty, DeepSeek, RubyLLM. One line to switch.

generate_with :anthropic, model: "claude-sonnet-4-5"
generate_with :openai,    model: "gpt-4o-mini"
generate_with :gemini,    model: "gemini-2.5-flash"
generate_with :ollama,    model: "qwen3:8b"
generate_with :deepseek,  model: "deepseek-flash"
generate_with :ruby_llm,  model: "mistral-large-latest"

# or per environment, in config/active_agent.yml
docs/providers ->

[]Tools from your schema

Bounded, allowlisted read tools generated from a model, scoped to whoever is asking.

class TicketTools < ActiveAgent::SchemaTools
  model Ticket

  filterable :status, :assignee_id, :opened_on
  returns    :id, :subject, :status, :opened_on

  scope_by_policy   # TicketPolicy::Scope, per caller
end

TicketTools.tool_names
# => ["find_tickets", "count_tickets", "get_ticket"]
docs/tools ->

{}MCP servers

Remote url: or local command: servers. Native where the provider supports it, bridged everywhere else.

def research
  prompt(
    params[:question],
    mcps: [
      { name: "github",
        url: "https://api.githubcopilot.com/mcp/",
        authorization: ENV["GITHUB_TOKEN"] },
      { name: "files",
        command: "mcp-server-filesystem",
        args: [ Rails.root.to_s ] }
    ]
  )
end
docs/mcps ->

->Delegation

Another agent as a tool, under a declared schema and a cost and latency budget.

class ClassifierAgent < ApplicationAgent
  delegation :classify, description: "Classify a ticket" do
    string :body, required: true
    returns { string :category, enum: %w[billing bug account] }
  end

  def classify(body:) = prompt(message: body)
end

class TriageAgent < ApplicationAgent
  delegate_to ClassifierAgent, budget: { max_calls: 1 }
end
docs/delegation ->

=Evaluations as tests

Replay scenarios across models, score them, and get one fault and one fix per failure.

scenarios = ActiveAgent::Evals::ScenarioParser.scenarios(<<~TEXT)
  # Orders
  Where is my order 4821? | tools: find_tickets
  Cancel my subscription  | not_contains: I cannot
TEXT

report = ActiveAgent::Evals::Runner.new(
  scenarios: scenarios,
  models: ActiveAgent::Evals::ModelSpec.parse_all(
    %w[claude-sonnet-4-5 ollama/qwen3:8b], default_provider: "anthropic"),
  replay: ->(scenario, spec) { SupportAgent.evaluate(scenario.prompt, model: spec.model, provider: spec.provider) }
).call

report.to_html   # or to_markdown, to_json, or publish it to the dashboard
docs/evaluations ->
  • #Structured output: JSON schemas as views
  • ~Streaming with open, chunk and close callbacks
  • :Retries with backoff, from each provider SDK
  • +Callbacks before, after and around every generation
  • @current_user on every call, tool and sub-agent
  • vReleases: a digest of what the model is given, on every trace
  • *Embeddings for search and RAG
  • ?A mock provider for tests, no keys needed
  • TToken usage and cost on every response
The dashboard · actionagent

See inside every agent decision.

Mount the engine in your app. Every generation becomes a trace, every conversation an interaction, every scenario suite a run you can compare against the last one. MIT, on your own database, from the first generation.

$ bundle add actionagent
$ rails generate action_agent:install && rails db:migrate

# config/active_agent.yml
telemetry:
  enabled: true
  local_storage: true      # open /activeagents
-> Traces All · Errors · 30s
> -> tr_9f2e41 SupportAgent#reply 2.4s~$0.0182v7 200
> -> tr_9f2e3c TriageAgent#triage 5.8s~$0.0410v7 200
> -> tr_9f2e2a ResearchAgent#research 12.1s~$0.1120v6 429
> -> tr_9f2e19 TranslationAgent#translate 0.9s~$0.0021v3 200
Traces. Every generation, with its prompt, LLM, tool and thinking spans, token counts (input, output, cached, thinking), provider, model and the release of the agent that produced it.
# Metrics 1h · 24h · 7d · 15 min buckets
Requests 12.8K +23%
Latency 847ms p95 2.1s
Error rate 0.4% 51 errors
Tokens 2.4M ↓1.8M ↑0.6M
Cost $42.18 $0.0033/req
Requests, stacked by agentdeploy and incident markers
v7 SupportAgent 429 · ResearchAgent
Metrics. The on-call view: golden signals, six time series over 1h, 24h or 7d, and the top agents, models, actions, tools and error types. A deploy marker names the version that moved the line.
= Evaluations · Support suite · run #12 +3 passed vs #11
Scenario claude-sonnet-4-5 judge's pick ollama/qwen3:8b
Passed 14/16 · 88% 11/16 · 69%
Cost per run ~$0.0243 · judge ~$0.0015 ~$0.0000
Where is my order 4821? find_tickets [+] 0.96 find_tickets [+] 0.81 find_tickets
Refund my last invoice refund [!] 0.42 missing_tool [!] 0.31 missing_tool
Cancel my subscription [+] 0.91 [!] 0.55 ungrounded_answer
What to fix missing_tool · 1 scenario · both models refund served by Stripe
Both models asked for a refund tool that this agent has available but not enabled. Enable Stripe for SupportAgent ->
Evaluations. Paste scenarios, pick the models to compare, and get a pass rate per model, a fault per failure and the fix each one calls for. Runs are pinned to the agent release they scored, so a stale suite says so.
<> Interactions ctx_4b1
userWhere is my order 4821?
agentfind_tickets order: 4821
tool1 result · shipped · 212ms
agentOrder 4821 shipped yesterday and arrives Thursday.
Context41.7k / 200k · 21%
messages tool results instructions schemas MCP
Interactions. The conversation behind each trace, with the context window as the model saw it.
> Run Agent SalesAgent
userHow did EMEA do this quarter? pipeline.csv
Revenue$3.36M
Deals113
Win rate31%
Break down by regionBook a follow-up
Message SalesAgent…Run ⌘⏎
Run Agent. Test an agent as a user would: a pinned conversation, editable context, file attachments, and generative UI the model renders back as cards, charts, forms and choices.
{} MCP server /activeagents/mcp
$ claude mcp add activeagents \
   https://app.example.com/activeagents/mcp
tools/list
run_support_agent · run an agent
find_tickets · schema tool, scoped to the key's caller
evaluations_run · replay a suite
traces_search · read failing traces
Agents as tools. The dashboard is itself an MCP server. Your coding harness can run agents, read records through schema tools, run evaluations and pull traces, all as the API key's caller.
[] Integrations Settings
  • 01GitHub connected · acme/support-app[+]
  • 02Checkout sandbox booted · mainready
  • 03Claude Code session · sonnetrunning
  • 04Suite re-run against the sandboxqueued
  • 05Diff: +41 -6 · instructions.mdreview
Integrations. Connect GitHub and an Anthropic or OpenAI key. A repository checkout boots as a sandbox, Claude Code or Codex works in it, and an evaluation runs against the checkout before anything merges.
+ Also on the sidebar
  • @Agents. Scorecards, versions, releases and the pooled eval pass rate per agent.
  • []Tools and MCP Services. The roster each agent is offered, with the server that serves each tool.
  • [>]Session Replay. Watch an agent-driven session step by step with its trace log beside it.
  • >Ask ActiveAgents. Ask which scenarios failed and why, with links to the evidence. Development and test only.
The platform · activeagents.ai

Same engine. Run for you.

Production observability for Rails AI agents, without hosting it yourself. Point telemetry at the platform and the team gets the dashboard: workspaces, retention, quotas and published evaluation reports. Every workspace starts with a free trial.

config/active_agent.ymldevelopment:
  telemetry:
    enabled: true
    # the mounted engine, in your own app
    local_storage: true

production:
  telemetry:
    enabled: true
    endpoint: https://api.activeagents.ai/v1/traces
    api_key: <%= ENV["ACTIVEAGENTS_API_KEY"] %>
    service_name: support-app
    # prompts and completions stay home unless you say so
    capture_bodies: false
  • @
    Workspaces and seats

    Invite the team. Every agent, trace, run and evaluation is scoped to the workspace, with provider keys of its own.

  • ->
    Hosted ingestion

    Traces post to /v1/traces and evaluation reports to /v1/evaluations, the same wire format a self-hosted mount accepts.

  • =
    Evaluation reports from CI

    A suite that ran in CI publishes its report into the workspace, pinned to the release it scored. The publisher checks the collector before a run is paid for.

  • #
    Retention and quotas by plan

    Trace retention from 3 to 400 days. Monthly trace and execution quotas that say when you are near them, with add-ons past them.

  • {}
    Your agents as an MCP server

    One authenticated endpoint at activeagents.ai/mcp serves your agents, schema tools, evaluations and traces to any MCP client.

  • []
    Managed sandboxes

    Checkout sandboxes and code sessions run on platform infrastructure, so nobody has to keep a laptop open to evaluate a branch.

Prefer to self-host? Everything above the accounts layer is MIT. Mount actionagent on your own domain and traces never leave your database. Self-hosted guide

Pricing

The gems are free. The platform scales with you.

activeagent, actionagent, solid_agent and activeagents-telemetry are MIT licensed. Self-host the dashboard with no seat or trace limits, or let the platform run it.

Free

trial

Enough to evaluate the platform. Not enough to run production on.

$0forever
  • +1 seat, 1 workspace
  • +250 traces a month, 3-day retention
  • +25 managed agent executions a month
  • +Traces, metrics, interactions and evaluations
  • +Community support on GitHub and Discord
Start free

Enterprise

annual

For organizations with compliance, scale and support requirements.

$269/ month

or $2,690 / year, billed annually

  • +Unlimited seats and workspaces
  • +500,000+ traces a month, 400-day retention
  • +Unlimited managed agent executions
  • +SSO / SAML, SOC 2, HIPAA, private VPC
  • +Multi-app and embedded licensing
  • +Dedicated Slack channel, 4-hour SLA
Contact sales

Compare self-hosted and platform plans

+
Self-hosted Free Pro Enterprise
Dashboard
Traces, metrics, interactions, evaluations[+][+][+][+]
Agent builder, Run Agent workbench, versions[+][+][+][+]
Agents as an MCP server[+][+][+][+]
Where traces liveYour databaseYour workspaceYour workspaceYour workspace or private VPC
Limits
Traces a monthUnlimited25025,000500,000+
RetentionYours to set3 days14 days400 days
Managed agent executions a monthUnmetered2510,000Unlimited
Seats / workspacesYours to set1 / 15 / 1Unlimited
Managed sandboxes and code sessionsLocal backend[ ][+][+]
Security and support
SSO / SAML and RBACYour auth[ ][ ][+]
SOC 2 Type II and HIPAAYour controls[ ][ ][+]
SupportCommunityCommunityEmail, 48hDedicated Slack, 4h SLA
Add-ons
Additional traces——$2.00 / 1k$1.50 / 1k
Additional executions——$0.01 each$0.005 each
Additional seats / workspaces——$19 seat / $49 workspaceIncluded
Services

Need a hand shipping?

The people who build the framework can build your first agents with you, review the ones you have, or take a project end to end.

[]Workshops

$2,500 half day · $4,500 full day

  • Build your first agent in a day, hands-on
  • Architecture patterns that hold up in production
  • Evaluations and tracing from the first commit
  • Code review and Q&A, custom curriculum available
Book a workshop

@Advisory

From $3,000 a month

  • A dedicated AI architect on your team
  • Strategy, code reviews and PR feedback
  • Async Slack access
  • Cancel any time with 30 days notice
Schedule a call

>Development

$250 an hour, or fixed price

  • We build the agents, you own the code
  • Integrated with your existing Rails app
  • Tuned, evaluated and deployed to production
  • A statement of work with clear milestones
Get in touch
FAQ

Questions people ask first

What is free and what is paid?

+

Every gem is MIT licensed: the framework, every provider, the dashboard engine, persistence and the telemetry adapters. Mount the dashboard in your own app and there are no seat or trace limits. The platform at activeagents.ai hosts that same dashboard for you, and its plans pay for ingestion, retention, quotas, managed sandboxes and support. Every workspace starts with a free trial.

How do I install it?

+

Add activeagent and run rails generate active_agent:install for the framework. Add actionagent and run rails generate action_agent:install followed by rails db:migrate for the dashboard; it mounts at /activeagents. Turn on telemetry.local_storage in config/active_agent.yml and every generation appears as a trace.

Which providers are supported?

+

OpenAI (Chat Completions and the Responses API), Anthropic, Google Gemini, AWS Bedrock, Azure OpenAI, Ollama, OpenRouter, Requesty, DeepSeek, and RubyLLM, which brings its own registry of models. A mock provider runs the whole pipeline in tests with no keys. Switch providers with one line.

Does MCP work with every provider?

+

Yes. Every provider accepts mcps:. Anthropic and the OpenAI Responses API run remote servers natively; for the others, and for every local command: server, the framework connects, lists the tools and exposes them as functions. Tool lists are cached, so a warm cache only connects when the model calls a tool.

How does the platform get my data?

+

Your app posts trace JSON to https://api.activeagents.ai/v1/traces, authenticated with a workspace API key, from one YAML block. Prompts, completions and tool arguments are sent only when capture_bodies is on; it is off by default. Error messages are truncated and backtraces are never sent. The wire format is open and the same one a self-hosted mount accepts.

Can I self-host the dashboard for a team?

+

Yes. The engine is built to be mounted in a shared Rails app: it resolves the signed-in user through your own session, scopes agents through a lambda you provide, accepts traces from other apps at <mount>/api/traces, and can run multi-tenant with an account per tenant. The self-hosted guide covers authentication, ingest keys and retention jobs.

What do evaluations actually test?

+

A scenario suite is a list of user messages with expectations: the tools a passing answer should call, text it must or must not contain. Each scenario replays against each candidate model, is scored by rules and optionally an LLM judge, and any failure is assigned exactly one fault with a recommendation. A run can also sample an agent's recent production generations instead of replaying. Runs are pinned to the agent release they scored, so a result from before the last instruction change reads as stale.

Does it work with an existing Rails app?

+

That is the point. Agents live in app/agents beside your controllers, prompt templates live in app/views, generations run inline or on Active Job, and current_user flows from your authentication into every callback, tool and sub-agent so Pundit or CanCanCan authorize an agent the way they authorize a controller. Rails 7.2 through 8.1 and Ruby 3.2 and up.