What is RAGfly?+−
A RAG service for AI agents: it turns any document corpus —thousands or tens of thousands, scans included— into a secure, multi-tenant retrieval base that's ready for production, without building or maintaining a RAG pipeline.
Who is it for?+−
For developers, consultancies and integrators building AI agents on private documents —sometimes for many different clients— who don't want to build or maintain a RAG pipeline at scale.
How much does it cost?+−
Plans in USD: Free $0, Starter $39/month, Growth $149/month, Scale $590/month and Enterprise by agreement. Each plan shows active corpus, Simple Retrieval, Augmented RAG, Agentic RAG and entities; paid plans support manual add-ons and have no automatic overage. MCP, REST, CLI and SDK are available on every plan.
How does it work?+−
You point RAGfly at your documents directory; the system scans, vectorizes and indexes them automatically. Then your agent retrieves by meaning and gets answers with citations to the source via MCP, REST or CLI, isolated by context and with the data handling that matches the deployment mode.
How does an AI agent consume RAGfly?+−
Through a remote MCP server (SSE or HTTP) or the RAGfly Desktop CLI, authenticating with a JWT or an API Key. The catalog of operations is published at `https://ragfly.ai/agents.json` and `https://ragfly.ai/llms-full.txt`.
What does RAGfly process, and where does my data end up?+−
RAGfly processes document information on the way in and relevant chunks on the way out. In Cloud it processes the original file but does not retain it after ingestion; chunks, embeddings and metadata persist. With Desktop, the original and its extraction are processed inside your network, although derivatives may be synchronized. On sovereign plans, the database and the derivatives at rest belong to the customer; RAGfly accesses the chunks it needs to retrieve and deliver context.
How is one customer's data isolated from another's?+−
Multi-tenant by default: each customer lives in a corpus isolated at the database level, with a Groups → Entities → Areas structure. A consultancy or integrator can serve dozens of customers from a single platform without a single record crossing over.
Can I use my own vector database and my own model?+−
It depends on the plan. Growth enables parsing-BYOK, embedding-BYOK and agentic-BYOK; Scale adds a customer-owned Supabase vector project; Enterprise supports customer LLMs and compatible vector databases.
Which AI models does it work with — Claude, GPT, Gemini?+−
RAGfly is model-agnostic. It delivers retrieval with citations and permissions; your agent runs on whichever LLM you prefer —Claude, GPT, Gemini or another— and consumes RAGfly via MCP, REST or CLI.
Which formats are supported? Does it work with scanned PDFs (OCR)?+−
PDF, Word, Excel and plain text, including scanned documents thanks to first-rate OCR and table and layout understanding. The focus is text and complex documents, not audio or video.
Can I use it with regulated data (legal, healthcare, government)?+−
Yes. Desktop processes the original file and its extraction inside the customer's network. On sovereign plans, the customer owns the database where the derivatives persist. RAGfly still processes chunks on the way out to produce retrieval; a full in-situ runtime is the mode that keeps that processing inside the customer's infrastructure too.
How long does it take to go from a directory to an agent answering with citations?+−
You point RAGfly at a directory and, with automatic ingestion, vectorization and indexing, you can have an agent answering with citations in a short time, with no human help — and without building or maintaining a RAG pipeline at scale.