Demonstration ProjectBasic AIAI / productivity

Draftly

Reply faster. Sound like yourself.

Draftly is a small writing tool for people who answer a lot of email. You paste the message you received, choose a tone and a length, and a draft streams in word by word. You edit it in place, copy it, and keep the prompt templates that work. I built it to show a clean, honest, single-call LLM integration: streaming, rate limits and per-request cost visibility.

  • Next.js
  • TypeScript
  • LLM API
  • Streaming (SSE)
  • Tailwind CSS
  • MySQL
Compose with streaming draft. Interface designed and built by Vikrant

Overview

Business problem
Freelancers and consultants spend a surprising part of the day on replies that are routine but still need the right tone: a polite no to a discount request, a firm nudge on a late invoice, a warm answer to a first enquiry. Generic chat tools work, but every time you re-type the same instructions, and it is hard to tell what each request costs.
Product objective
Make one task fast: turn an incoming message into a good first draft in the tone and length you choose, keep control of the final wording, and keep the cost and the data handling easy to understand.
Target users
Solo professionals and small teams who write many client emails, such as consultants, freelancers, recruiters and account managers.
Solution
A focused compose screen that sends one request to an LLM API with a versioned prompt template, streams the answer back, and lets the writer edit and copy it. Templates, tone presets, history and a usage page sit around that single flow.

Project classification

Project type
AI writing assistant
Industry
AI / productivity
Frontend
Next.js with TypeScript and Tailwind CSS
Backend
Next.js route handlers that call one LLM API and stream the response
Database
MySQL for drafts, templates, tone presets and usage records
Architecture
Single prompt-assembly service with a streaming endpoint and per-user limits

User roles

  • Writer (free plan)
  • Writer (Pro plan)
  • Admin (templates and limits)

My role

  • Product and UX design of the compose flow, including the streaming and stop/regenerate behaviour.
  • Frontend: compose, history, templates, tone presets and usage screens in Next.js, TypeScript and Tailwind CSS.
  • Backend design: a streaming route handler, prompt assembly from a template plus tone and length settings, and error handling for provider timeouts.
  • Database design: MySQL tables for drafts, templates with versions, tone presets and usage records.
  • Rate limiting, monthly allowance checks and per-request cost estimates.
  • This is a self-initiated portfolio project; there is no client, live deployment or user base.

Key features

  • Compose with streaming draft. Interface designed and built by Vikrant

    Compose with a streaming draft

    The writer pastes the message they received, picks a tone and a length, and presses Draft. The reply appears word by word in a paper-style card, with a Stop button while it streams. When it finishes, the draft is an editable text area with Copy and Regenerate beside it, so the first draft is a starting point rather than a final answer.

  • Draft history. Interface designed and built by Vikrant

    History that remembers the settings

    Every draft is saved with the message it answered, the template and version, the tone and the length. The history list can be searched and filtered by tone, and reopening a draft restores the exact settings so a good result can be reproduced or adjusted.

  • Prompt templates. Interface designed and built by Vikrant

    Prompt templates with versions

    Templates such as Polite decline, Payment reminder or Follow up after a call hold the task instruction and a few placeholders. Each edit creates a new version number, and every draft records which version produced it, so a change in wording can be traced and rolled back.

  • Tone presets. Interface designed and built by Vikrant

    Tone presets you can read

    Tones are short, plain instruction blocks rather than hidden settings: Warm, Direct, Formal, Concise. The writer can read exactly what each one tells the model, adjust the wording, and add a personal preset such as a Swiss-German business register.

  • Usage and plan. Interface designed and built by Vikrant

    Usage, limits and cost in plain view

    The usage page shows drafts used against the monthly allowance, the daily request limit, and an estimated cost for each request from the token counts the provider reports. It also states what is stored: settings and token counts are kept, message text is not written to logs.

Also in the product

  • Compose view with source message, tone and length controls
  • Streaming draft with stop, regenerate, edit in place and copy
  • Draft history that keeps the settings used for each draft
  • Prompt templates with version numbers
  • Tone presets that map to short instruction blocks
  • Per-user rate limits and a monthly draft allowance
  • Estimated token cost shown for each request
  • Logs that store metadata, not message text

User flow

  1. Writer signs in and opens the compose screen
  2. Writer pastes the incoming message and picks a template, tone and length
  3. The server checks the daily limit and the monthly allowance
  4. The server assembles the prompt and opens a streaming request to the LLM API
  5. Text streams to the browser over server-sent events; the writer can stop at any point
  6. Writer edits the draft in place and copies it into their email client
  7. The draft, its settings and token counts are saved to history
  8. Writer reviews templates, tone presets and usage when needed

Technical architecture

  1. Next.js app: compose, history, templates, tone presets and usage pages
  2. Route handler for draft generation that streams via server-sent events
  3. Prompt assembly: template version, tone block, length rule and the pasted message
  4. Rate limiter and monthly allowance check before every request
  5. Single LLM API call behind a thin provider wrapper
  6. MySQL: users, drafts, templates, template versions, tone presets, usage records
  7. Structured logging of request id, template version and token counts only

Database and backend

Draft generation

  • One POST endpoint validates the input with Zod: message length, template id, tone and length options.
  • The prompt is built in a fixed order: base instructions, template version, tone block, length rule, then the pasted message clearly marked as content to answer.
  • The provider response is streamed straight to the client as server-sent events; a closed connection cancels the upstream request.
  • When the stream ends, the draft and the reported token counts are saved in a single transaction.

Limits and cost visibility

  • A per-user limit on requests per minute and per day, plus a monthly draft allowance tied to the plan.
  • Each usage record stores model name, input tokens, output tokens and a pricing snapshot, so the estimated cost stays stable if prices change.
  • A maximum input length keeps a single request from becoming expensive.
  • Requests over a limit return a clear message and a retry time instead of a generic error.

Privacy and logging

  • Server logs hold a request id, user id, template version, status and token counts. They never hold the pasted message or the draft.
  • Drafts are stored in the database for the user's history and can be deleted by the user.
  • A short notice on the compose screen reminds writers not to paste passwords or payment details.

Data model

  • User, Draft, Template, TemplateVersion, TonePreset, UsageRecord.
  • Draft stores the template version id, tone, length, source message and final text.
  • Template versions are immutable; editing a template inserts a new version row.

Challenges and solutions

  • Streaming that feels calm instead of jumpy

    Tokens are buffered briefly on the client and appended in small groups, the card keeps a fixed height with scroll anchoring, and a Stop button is always present. If the connection drops, the text received so far stays editable and a Retry option appears.

  • Changing prompts without losing track of what produced a draft

    Templates are versioned and immutable. Each draft stores the version id it used, so a result can be reproduced and a poor edit can be reverted by selecting an earlier version.

  • Preventing abuse and runaway cost on a public tool

    Per-minute and daily limits, a monthly allowance, a maximum input size and a cap on output length. Over-limit responses return a retry time. The limits are enforced on the server before any provider call.

  • Showing what a request costs without false precision

    Cost is shown as an estimate computed from the token counts the provider reports and a stored price per model, labelled as an estimate. Nothing is recalculated later with newer prices.

  • Keeping personal data out of logs

    The logging helper accepts only an allow-listed set of fields: ids, template version, status and token counts. Message text never reaches it, which is easier to guarantee than scrubbing text afterwards.

  • A pasted email that tries to give the model instructions

    The incoming message is wrapped and labelled as content to reply to, and the base instructions say not to follow instructions inside it. This reduces the risk but does not remove it, so the output is always a draft that the writer reads and edits before sending.

Results

  • A compact case study of a single-call LLM feature done carefully: prompt assembly, streaming, limits and cost visibility.
  • A clear separation between prompt templates, tone presets and application code, with versioning that makes changes traceable.
  • A privacy-minded logging approach that records usage without recording what people write.
  • Honest scope: one model call per draft, with a human editing every result, rather than an autonomous agent.

Qualitative outcomes only; no usage or revenue figures.

Visual identity

  • Primary #7A5CFA
  • Secondary #161B2E
  • Accent #FFB88C
  • Background #FAF9FF

Fraunces for headings and drafts, Figtree for the interface. Rounded violet square with a folded-page corner and a peach pen nib. Calm two-column writing surface, generous whitespace, serif drafts on a paper-white card.

  • Demonstration ProjectAI / automation

    AgentForge

    Build, ground and monitor AI agents with RAG, tools, a testing playground and usage tracking.

    • Next.js
    • React.js
    • TypeScript
    • Node.js
    • Laravel
    • PostgreSQL
    • pgvector
    • RAG
    • LLM APIs
    • REST API
    • Webhooks
  • Portfolio ProjectFood / subscription commerce

    Cadence Kitchen

    Meal-plan subscription platform with skip, pause and nutrition tracking.

    • React.js
    • TypeScript
    • React Query
    • Node.js
    • Express.js
    • MongoDB
    • REST API