Conversational AI Grounded in Your Own Content
Configure AI Assistants powered by the Claude model family, back them with vector-based semantic search over your content, and let RAG ground every answer in what's actually in your system — with persistent chat threads instead of a stateless chatbot widget.
AI That Knows Your Content, Not Just Language
Choose the right Claude model per use case
Claude Opus for complex reasoning and creative generation, Claude Sonnet for general conversation and recommendations, Claude Haiku for fast, high-volume FAQ and lookup traffic.
Custom system prompts and behavior
Configure each assistant's system prompt, temperature, and max token limit, and enable tool integration so it can search content, query data, or call other functions.
Vector-based semantic search
Content is embedded into vector indexes so search understands meaning, not just keyword overlap — with hybrid search combining semantic and keyword matching when that's the better fit.
RAG grounds responses in real content
Retrieval Augmented Generation pulls relevant content via vector search before generating a response, so answers stay factually grounded, current, and attributable to a source.
Persistent chat threads
Conversations are managed as user-specific threads that preserve context and message history, so a customer's follow-up question doesn't need to restate what they already said.
Tool integration and function calling
Assistants can call configured tools — vector search, database queries, API integrations — to extend what a conversation can look up or act on, beyond the model's own knowledge.
From Content Index to Grounded Conversation
Configure an assistant
Set a name, system prompt, model, temperature, and max token limit, and decide which tools it's allowed to call.
Index your content
Build a vector index over your content so the assistant has something concrete to retrieve from when it answers a question.
Converse
Threads keep conversation context and history per user, across sessions, so multi-turn interactions stay coherent.
Ground every response
RAG retrieves relevant content via vector search first, then generates a response grounded in it — accurate, current, and source-attributed.
See AI Assistants in Action
Configuration, search, and conversation, all in one place.
Configure a Claude-powered assistant's behavior without writing infrastructure code.
Test semantic search results before wiring an assistant to the index.
A persistent thread keeps conversation context across turns and sessions.
Enable the tools an assistant is allowed to call, per assistant.
AI Personalization Without a Separate AI Stack
Claude-powered assistants, vector search, and RAG — pre-integrated with the customer and content data already in Stack9, instead of a chatbot service, a search engine, and a recommendation system bolted together.