Skip to main content

Conversational AI Grounded in Your Own Content

Configure AI Assistants powered by the Claude model family, back them with vector-based semantic search over your content, and let RAG ground every answer in what's actually in your system — with persistent chat threads instead of a stateless chatbot widget.

AI That Knows Your Content, Not Just Language

Choose the right Claude model per use case

Claude Opus for complex reasoning and creative generation, Claude Sonnet for general conversation and recommendations, Claude Haiku for fast, high-volume FAQ and lookup traffic.

Custom system prompts and behavior

Configure each assistant's system prompt, temperature, and max token limit, and enable tool integration so it can search content, query data, or call other functions.

Vector-based semantic search

Content is embedded into vector indexes so search understands meaning, not just keyword overlap — with hybrid search combining semantic and keyword matching when that's the better fit.

RAG grounds responses in real content

Retrieval Augmented Generation pulls relevant content via vector search before generating a response, so answers stay factually grounded, current, and attributable to a source.

Persistent chat threads

Conversations are managed as user-specific threads that preserve context and message history, so a customer's follow-up question doesn't need to restate what they already said.

Tool integration and function calling

Assistants can call configured tools — vector search, database queries, API integrations — to extend what a conversation can look up or act on, beyond the model's own knowledge.

From Content Index to Grounded Conversation

1

Configure an assistant

Set a name, system prompt, model, temperature, and max token limit, and decide which tools it's allowed to call.

2

Index your content

Build a vector index over your content so the assistant has something concrete to retrieve from when it answers a question.

3

Converse

Threads keep conversation context and history per user, across sessions, so multi-turn interactions stay coherent.

4

Ground every response

RAG retrieves relevant content via vector search first, then generates a response grounded in it — accurate, current, and source-attributed.

See AI Assistants in Action

Configuration, search, and conversation, all in one place.

Stack9 AI assistant configuration screen showing system prompt, model picker, and temperature slider

Configure a Claude-powered assistant's behavior without writing infrastructure code.

Stack9 vector index management screen showing indexed documents and semantic search test results

Test semantic search results before wiring an assistant to the index.

Stack9 chat thread interface showing a multi-turn conversation with preserved context

A persistent thread keeps conversation context across turns and sessions.

Stack9 tool integration configuration panel showing available functions an assistant can call

Enable the tools an assistant is allowed to call, per assistant.

AI Personalization Without a Separate AI Stack

Claude-powered assistants, vector search, and RAG — pre-integrated with the customer and content data already in Stack9, instead of a chatbot service, a search engine, and a recommendation system bolted together.