AI settings
The assistant suggests replies, rewrites text and can answer customers on its own. The project brings its own provider and key.
The section is “Settings → Operator tools → AI Integration”; the owner and admins can change it. The AI usage page with spending sits next to it.
Connection
- Provider: OpenAI, Anthropic, Google Gemini, OpenRouter, GitHub Models or Custom (self-hosted) — for the latter you give the API base URL of an OpenAI-compatible server.
- The API key is stored encrypted and never shown back. Leave the field empty to keep the saved key.
- Test Connection sends a short request to check the key and the model.
Models and spending
- Main model — for reply suggestions, rephrasing and auto-replies. The list shows the price per 1M tokens next to each model; you can also enter a model manually.
- Fast model — for AI tags and department suggestions. When unset, the main model is used.
- Custom pricing (USD per 1M tokens) — Input and Output prices for a custom server and OpenRouter, used to calculate cost.
- Monthly budget — a soft cap in US dollars: the assistant keeps working, but the AI usage page warns when spending gets close to the budget or goes over it.
How the assistant replies
- Reply tone — Business (concise and formal), Concise (1–2 sentences), Friendly (warm, light emoji allowed) or System Prompt only (no presets).
- System Prompt — the global instruction for every AI call.
- Reply length cap — 100 to 1500 tokens per reply, with an estimated cost per reply next to it.
What the assistant does
- AI Reply Suggestion — the “AI Suggest” button above the reply field: a draft reply lands in the composer. When off, or not in the plan, the button isn't shown.
- Auto-reply (always) — AI answers Telegram customers by itself. Questions about money and refunds, complaints and requests for a human go straight to an operator; it answers only when confident, and no more than three times per ticket.
- Auto-reply outside hours — the switch is saved but has no effect right now.
- AI tags — the “AI tags” action in the ticket’s “…” menu: AI suggests tags from the conversation, and the operator applies or dismisses them. When off, the menu has no such item.
- Style Learning — suggestions copy the operator's style once they have the “Minimum closed tickets to activate”.
Rephrasing the text in the composer (🪄) and the department suggestion (“AI routing” in the ticket’s “…” menu) work whenever AI is connected, with no separate switch.
Knowledge base and ticket memory
An operator's reply suggestion pulls in relevant published knowledge base articles, public and internal alike, and answers from closed tickets: the operator reviews it before sending.
An auto-reply that goes to the customer without an operator takes only published articles marked Public from the knowledge base and works from the current conversation. It doesn't use the memory of past tickets: those are other customers' conversations.
- A published article is indexed in the background every time it's saved; when an operator closes a ticket, AI distils the question and answer from it into memory.
- Search is semantic when the provider supports embeddings (OpenAI, OpenRouter, GitHub Models, Gemini, a custom server). With Anthropic, keyword search is used.
- To index what existed before AI was connected, call
POST /api/w/{workspace_id}/settings/ai/rag/backfill— it goes through every published article and the 500 most recent closed tickets. There is no button for this in the interface.
AI usage
Every AI call is logged with its tokens and cost at the model's price. The AI usage page shows totals for the period (calls, tokens, total), spend by day, feature and model, a 30-day forecast and spend against the budget.
✨ AI mark on messages
When an operator sends a reply that started from an AI suggestion or was rephrased by AI (even after edits), the message keeps an AI mark. Only staff see it; the customer doesn't.

