Demonstration ProjectAI / automation

AgentForge

Build agents that do the work.

AgentForge is a workspace where a business creates AI agents, grounds them in its own documents and website, connects them to tools and webhooks, tests every change in a playground, and watches conversations, token usage and cost once they are live. I designed and built it to show how I architect an LLM platform end to end.

  • Next.js
  • React.js
  • TypeScript
  • Node.js
  • Laravel
  • PostgreSQL
  • pgvector
  • RAG
  • LLM APIs
  • REST API
  • Webhooks
Overview dashboard. Interface designed and built by Vikrant

Overview

Business problem
Teams that try to ship an AI assistant usually end up with a prompt in a script, no way to see why an answer was given, no control over what the model may do, and a surprise bill at the end of the month. Support, sales and operations each need a different agent, and nobody can test a change before customers see it.
Product objective
Give a non-research team one place to define an agent, ground it in approved content, restrict what it can do, test it against real questions, and measure what it costs and how often it resolves the conversation.
Target users
Small and mid-size businesses running customer support, lead qualification, bookings and internal help desks, plus the analysts and owners who have to answer for cost and quality.
Solution
A versioned agent builder, a retrieval pipeline with citations, a tool layer with guardrails and confirmation rules, a playground that exposes the full trace of every run, and metering that records tokens per workspace, agent and model.

Project classification

Project type
AI SaaS platform
Industry
AI / automation
Frontend
Next.js with React and TypeScript
Backend
Node.js orchestration service plus a Laravel admin and billing API
Database
PostgreSQL with a pgvector index
Architecture
Orchestrator, retriever and ingestion queue behind a REST API

User roles

  • Workspace Owner
  • Agent Builder
  • Analyst / Viewer
  • End user (chat widget)
  • Platform Admin

My role

  • System architecture: the split between the Next.js app, the Node.js orchestrator, the ingestion queue and the Laravel billing API.
  • Backend design: orchestrator flow, tool-call contract, guardrail checks and streaming responses.
  • Database design: workspace-scoped schema, agent versioning, chunk and embedding tables, usage records.
  • RAG pipeline design: chunking strategy, embedding, retrieval, reranking and citation format.
  • Frontend: builder, playground, knowledge base and usage screens in Next.js and TypeScript.
  • Integrations: webhooks, tool definitions with JSON Schema, LLM provider abstraction.
  • This is a self-initiated concept; there is no client, live deployment or user base.

Key features

  • Overview dashboard. Interface designed and built by Vikrant

    One overview for volume, quality and cost

    The dashboard puts conversations, active agents, resolution rate, average response time, token usage, estimated cost, leads captured and bookings made side by side, so the owner and the analyst read the same numbers. Recent conversations sit underneath, each tagged with the agent that handled it and whether it resolved or escalated.

  • Agents list. Interface designed and built by Vikrant

    A fleet of purpose-built agents

    Support, sales qualification, booking, restaurant reservations, real-estate leads and an internal knowledge assistant live in one list with status, model, volume and resolution. Each agent has its own instructions, knowledge, tools and budget, and drafts stay separate from the published version.

  • Agent builder. Interface designed and built by Vikrant

    Agent builder with versioned instructions

    A split-pane builder keeps the system instructions in a monospace editor on one side and the knowledge sources, tools, model, temperature, memory and guardrails on the other. Every publish creates an agent version, so a bad prompt change can be rolled back in one step.

  • Playground with citations and trace. Interface designed and built by Vikrant

    Playground with citations and a trace

    Before publishing, a builder asks the agent a real question and sees the answer with numbered citations, the chunks that were retrieved and their scores, and a trace of each step: guardrail checks, retrieval, reranking, tool calls and generation, with latency and token counts.

  • Knowledge base. Interface designed and built by Vikrant

    Knowledge base from documents and websites

    PDFs, Markdown files and site crawls are ingested through a queue: parsed, split into chunks, embedded and indexed. The knowledge screen shows each source with its status, chunk count and last sync, and surfaces failed or stale sources before they hurt answer quality.

  • Conversation transcript. Interface designed and built by Vikrant

    Conversation transcripts you can audit

    Every conversation keeps the full message history, the tools that were called with their arguments, the sources that were cited and the token cost. A reviewer can see why an agent said something and flag the answer for a knowledge or instruction fix.

  • Usage and cost. Interface designed and built by Vikrant

    Usage and cost under control

    Token usage is metered per request and rolled up by day, agent and model. The usage screen shows month-to-date spend against a workspace budget, which agents consume the most, and the split between input and output tokens, so cost surprises show up early.

Also in the product

  • Agent builder with versioned system instructions
  • Knowledge base from documents and website crawls (RAG)
  • Testing playground with source citations and a step-by-step trace
  • Tools and webhooks with confirmation rules
  • Guardrails, memory and model settings per agent
  • Conversation history and transcripts
  • Token metering, cost estimates and budget caps

User flow

  1. Owner creates a workspace and invites builders and viewers
  2. Builder creates an agent and writes its system instructions
  3. Builder uploads documents or adds a website to crawl
  4. Ingestion parses, chunks, embeds and indexes the content
  5. Builder enables tools, sets guardrails, model and temperature
  6. Builder tests questions in the playground and reads the trace
  7. Builder publishes a version and embeds the chat widget or connects a webhook
  8. Customers talk to the agent; messages stream back with citations
  9. Analyst reviews conversations, resolution and escalations
  10. Owner tracks token usage and cost against the budget

Technical architecture

  1. Next.js app (builder, playground, dashboards) and embeddable chat widget
  2. REST API with workspace-scoped auth and request validation
  3. Orchestrator (Node.js): prompt assembly, tool calls, guardrails, streaming
  4. Retriever: query embedding, vector search (pgvector), reranking
  5. LLM API providers behind a single provider interface
  6. PostgreSQL with pgvector for agents, conversations, chunks and usage
  7. Ingestion queue and workers: parse, chunk, embed, index
  8. Outbound webhooks and tool endpoints; Laravel billing and admin API

Database and backend

RAG pipeline

  • Ingest: uploads and crawls create a Source and a Document row, then enqueue a job; the API returns immediately and the UI polls status.
  • Chunk: structure-aware splitting on headings and paragraphs with a token-size target and a small overlap, keeping the heading path with every chunk for citations.
  • Embed: batched calls to the embedding model; each chunk stores its vector, token count and a content hash so unchanged text is never re-embedded.
  • Retrieve: the question is embedded and searched in pgvector, filtered by workspace and agent knowledge scope, then merged with a keyword match for exact terms such as SKUs and policy names.
  • Rerank: the top candidates are re-scored against the question and trimmed to the few that fit the context budget; their source and heading path become the citations.

Orchestrator, tools and guardrails

  • Prompt assembly: system instructions, agent memory summary, retrieved chunks, tool definitions and the recent turns, in a fixed order with a per-section token budget.
  • Tool calls: tools are declared with JSON Schema; arguments are validated before execution and the result is validated again before it goes back to the model.
  • Confirmation rules: tools that change something (returns, bookings) can require explicit user confirmation or builder-defined conditions.
  • Input guardrails: prompt-injection and off-topic checks before generation; output guardrails check for leaked instructions, personal data patterns and missing citations.
  • Loop limits: a maximum number of tool iterations and a per-run timeout end runaway loops with a safe fallback answer and a handoff offer.

Token metering and billing

  • Every model call writes a UsageRecord with workspace, agent, version, model, input tokens, output tokens and a pricing snapshot.
  • Costs are estimated from the stored pricing table, not recalculated later, so history stays stable when prices change.
  • Workspace budgets raise warnings at a threshold and stop new conversations at a hard cap, with the owner notified.
  • A Laravel service owns plans, invoices and admin tooling and reads usage through an internal API.

Data model

  • Workspace, Agent, AgentVersion, KnowledgeSource, Document, Chunk (with embedding), Conversation, Message.
  • ToolDefinition, Integration, Webhook, Guardrail, UsageRecord.
  • Every row carries a workspace id; queries are scoped in one data-access layer, backed by Postgres row-level security.

Challenges and solutions

  • Keeping long conversations inside the context window

    Prompt assembly uses a token budget per section. Recent turns are kept verbatim, older turns are folded into a running summary stored as agent memory, and retrieved chunks are trimmed by rerank score rather than cut off mid-passage.

  • Retrieval that returns the right passage, not a plausible one

    Chunks follow document structure instead of fixed sizes and keep their heading path. Vector search is combined with keyword matching for exact terms, results are reranked, and the playground shows scores so builders can fix weak sources instead of guessing.

  • Letting an agent call tools without letting it do damage

    Tools are allow-listed per agent, arguments are validated against JSON Schema, write actions can require confirmation, and loops are capped. Every call is logged in the trace with its arguments and result.

  • Streaming answers while guardrails still apply

    Tokens stream to the widget over server-sent events. Input checks run before generation, and output checks run on a sliding window, so a violating answer is stopped and replaced with a safe message rather than shown in full and retracted.

  • Isolating one workspace's data from another's

    Workspace id is on every table, including chunks and embeddings. A single data-access layer sets the tenant on the connection, Postgres row-level security enforces it, and vector queries always include the workspace filter.

  • Metering tokens accurately across providers and retries

    Usage is recorded from the provider's reported token counts per call, keyed by request id so retries are not double-counted, with a pricing snapshot stored alongside each record.

Results

  • A complete case study of an LLM platform covering retrieval, orchestration, tool safety and cost visibility, not just a chat box.
  • A schema and service split that keeps tenant data separate and keeps billing out of the hot path.
  • A testing workflow where every change can be inspected step by step before it reaches customers.
  • Shows the product thinking behind AI features: grounding, limits, observability and budgets.

Qualitative outcomes only; no usage or revenue figures.

Visual identity

  • Primary #F25C19
  • Secondary #1A1B1F
  • Accent #6D5EF5
  • Background #FAFAF8

Inter Tight for the interface, JetBrains Mono for prompts, JSON and traces. Anvil-and-spark glyph with a plain wordmark. Neutral light workspace, split-pane builder, monospace prompt editor and trace panels.

  • Portfolio ProjectFood / subscription commerce

    Cadence Kitchen

    Meal-plan subscription platform with skip, pause and nutrition tracking.

    • React.js
    • TypeScript
    • React Query
    • Node.js
    • Express.js
    • MongoDB
    • REST API
  • Demonstration ProjectBasic AISmall business / e-commerce

    Pebble Chat

    An embeddable FAQ chat widget for a small shop that answers from the owner's own policies and hands anything uncertain over by email.

    • React.js
    • Node.js
    • Express.js
    • LLM API
    • REST API
    • SQLite