The AI assistant
Cushy's AI assistant: real model calls via any OpenAI-compatible backend, grounded every turn on a live digest of your org's real cost, inventory, topology graph and operations data, with RAG over prior Q&A and indexed insight documents, plus image attachments.
The assistant overlay makes real model calls — there are no canned answers. When no model backend is available the assistant says so honestly rather than fabricating.
Grounded on your real infrastructure
Every question is answered against a live, org-scoped data digest rebuilt each turn from your real platform state: actual spend (30-day total and trend, spend by provider, top services, daily figures, open anomalies, budget consumption), the synced inventory (counts by provider/type, health, regions, Terraform-managed share), the network topology graph (networks, subnets, gateways, cross-cloud connections and up to 25 relationship facts like which instance sits in which subnet and VPC), and operations state (open alerts, workspace apply history). Ask *"analyse my cost"* and the assistant answers with concrete numbers and trends — setup guidance appears only when the data genuinely doesn't exist yet. No cross-org data and no secrets ever reach the model.
On top of the digest, RAG retrieval over an in-database vector store (strictly org-scoped) grounds each answer on your organization's most similar prior questions & answers plus indexed insight documents — per-provider cost summaries, top-service and anomaly sentences, and topology graph facts, refreshed automatically after each cost ingestion and periodically thereafter. Every turn's Q&A is stored back for future retrieval. If embeddings are unavailable the assistant still answers, just without retrieval.
Image attachments
Attach up to 4 images per question (PNG, JPEG, WebP or GIF, ≤5MB each) with the 📎 button — a screenshot, an architecture diagram, an error dialog. They are validated server-side and forwarded to the model as multimodal content, and they are transient: sent with the turn, never stored. A turn may be image-only.
Failure modes are honest: no backend configured → 503, backend unreachable → 502 — never an invented answer. Requests are per-user rate-limited and audited (assistant.query). Streaming responses are a noted follow-up.
The deployer assistant
Inventory and Terraform carry a second, narrower assistant: ✦ Ask to deploy. Describe what you need in plain language — "a small web server in eu-west-1" — and it drafts the deployment: which blueprint, which connected account, which Terraform folder, and the variables it can infer.
It proposes; it never deploys. The draft opens in the ordinary Deploy a new resource wizard, which runs the same plan → policy → cost → approve → apply pipeline as a hand-built deploy. You always see the real Terraform diff, the policy result and the cost delta before anything is created.
- It can only pick a blueprint that is really in your catalog — a name it invents resolves to nothing and is reported as unmatched, never swapped for a lookalike.
- Anything it can't safely infer — a VPC id, an AMI, a password — is left blank and handed back as a field for you to fill.
- It needs the
mutate_infracapability; a read-only role is refused. - Every proposal is audited as
assistant.deploy_proposed.
A model that could call the cloud directly would be a second, unaudited mutation path. Keeping the assistant on the proposing side means an assistant-originated deploy and a hand-built one are the same deploy, through the same routes, with the same approvals.