RAG Service

Augmented retrievalfor your agents.

Simple, private, and precise. And each of your contexts, in its own space.
Works with the agents you already use — over MCP — and with the ones you build.
At scale and isolated per context, with privacy verifiable by deployment mode. Connect a directory and your agent answers with citations — without building or maintaining a RAG pipeline.
PREFLIGHT · THE REFUSAL

Verified answers, not plausible ones.

A verifier breaks the answer into claims and checks each one against its evidence before showing it. If something isn't backed, it doesn't ship.

You can only trust a system that knows how to say no.

Every claim that does pass the check comes with its citation, linking straight to the document and the page.

PREFLIGHT
  • corpus indexedPENDINGVERIFIED
  • permissions appliedPENDINGVERIFIED
  • citation resolved to document and pagePENDINGVERIFIED
  • evidence for this claimevidence for this claimPENDINGNO EVIDENCE
SYSTEM ANSWER

“I found no evidence in your documents to answer this with confidence.”

SIMPLE · BOARDING

You point at the directory. The pipeline does the rest.

You bring the directory —your local folder— and the system loads, extracts, analyzes, chunks and vectorizes on its own. No document ingestion to build, no upfront configuration and no integration project.

PIPELINE STATES
  1. CARGADO
  2. METADATA
  3. ESCANEADO
  4. CHUNKEADO
  5. VECTORIZADO

The climb is smooth; the cruise, instant.

Ingesting a large corpus runs on its own, unattended: you leave it going and move on. What you notice in every query isn't speed — it's that the answer is faithful to what you loaded.

Onboarding without red tape

If you start out alone, group, entity, area and profile are a single node that unfolds as your organization grows. The simplicity of day one, with the enterprise ceiling already built in.

INPUT · FORMATS

PDF, Word, Excel, PowerPoint, RTF, plain text and scans via OCR.

What we don't doWe don't process audio, video, or non-text images.

OUTPUT · INTERFACES

Web chat, MCP server, REST API /v1, Python SDK, TypeScript SDK and CLI.

On every planAll six, from a single configuration. RAGfly Desktop is separate, from Starter up.

PRIVATE · THE MANIFESTO

We never store your document in the cloud. And if you want, the document never leaves your network.

In Cloud the original is processed on the way in and not retained: what persists are the derivatives —chunks, embeddings and metadata—. With Desktop, the file and its extraction stay inside your network, and depending on the deployment encrypted derivatives may be synchronized. On sovereign plans, that database is yours.

ISOLATION

It's not that you can't see it. It's that it isn't on board.

The group and entity filter lives inside the search engine itself: vectors from another context are never even scored. It isn't a filter applied afterwards — they never entered the search. It holds on every plan, the free one included.

Uploading vectorized content is not uploading the document.

Extracted text goes up, in chunks, encrypted. The file doesn't. For example, if the document is signed, the signature doesn't go up.

MI-PASS

Group → Entity → Area

The group is the tenant; the entity separates organizations or contexts within the group; the area defines which document subtree each identity sees, with hierarchical inheritance — whoever sits above sees what their descendant areas hold.

With no area assigned, access is closed by default, never open.

The text of a forbidden document never reaches the model or the agent: it isn't hidden on screen, it doesn't enter the prompt.

YOUR CREDENTIALS

With parsing-BYOK and embedding-BYOK, no model under our account ever sees your document.

The five pipeline calls —reading, typing, summary, attributes and vectorization— run on your credentials, against Anthropic, OpenAI or Google, under your contract and your invoice.

We never retry with ours.

If your credential fails —quota, rate limit, provider down—, the operation fails and we tell you. A silent fallback would break exactly that promise.

PRIVACY FLOOR

Set by the plan, not by a checkbox.

Paid plans resolve only against model providers contractually bound not to train on the content sent — not a checkbox someone can leave in the wrong position.

Verified by an automated suite that mirrors the resolution across all 462 plan × skill combinations, with a positive control.

The free plan runs on low-cost models that don't include that clause: it's for evaluation, not for sensitive documents.

INSTRUMENTS
AES-256-GCM · application-level encryption with key hierarchyTLS 1.3 in transitAudit log of business operations
PRECISE · THE TRAIL

Retrieves by meaning and shows where it came from.

The answer doesn't end at the text: it ends at the exact chunk that backs it, with its page.

CITATIONS

Direct link to the document and the page

Every claim carries its citation. The trail runs from the answer to the chunk that backs it, and you can follow it.

RETRIEVAL

Hybrid search + rerank

Three branches —vector search over pgvector, lexical search over what the system understood from the document, and a dedicated branch for the summary— fused and reordered by a multilingual reranker.

If rerank isn't available, the degradation is visible, not silent.

EVALUATION

Precision measured, not declared

An evaluation suite with verifiable ground truth: the question is generated from the real chunk that answers it, so the correct answer is known in advance.

TRIAGE BY LEVEL
  1. 01 · retrieval
  2. 02 · generation
  3. 03 · citation
  4. 04 · numeric accuracy

Improvement is targeted, not guesswork.

CONFIGURATION

Configuration expands the context.

Everything you configure becomes context your agent receives — editable from the screen, with no deploy and no developer. It's what separates RAGfly from a vector database with a chat on top.

SYSTEM PROMPTS01

Seven layers, five editable

Product → application → function → your business hierarchy (group, entity, area, profile) → capability catalog. You write «this group is a construction company; its documents are usually works contracts» and from the next message on, everyone answers through that lens.

A change in a text box, new behavior instantly.

TYPOLOGY02

Types and attributes

The pipeline doesn't chunk blindly: it types each document and extracts only the attributes that apply to its type —amount, counterparty, due date—.

That's what enables exact answers: «how many invoices from May?», «the contract with the highest amount».

ENRICHMENT03

Context in every vector

Before vectorizing, each chunk is prefixed with the whole document's record: name, location, date, summary and attributes. Enrichment, not just chunking.

It's measured: we tried removing it and retrieval got worse.

The summary also competes as a candidate of its own in the search.

PROFILES · THE CREW

You don't give your agent access to the documents. You give it a profile.

The agent signs in as «Finance Analyst» and sees exactly what that profile would see — under the same access rules as your people.

Service identity

01

A profile is an identity with no personal data, designed for agents. It inherits the same machinery as a person: area with hierarchical inheritance, associated locations, roles, RBAC and identity prompts. No master keys.

Frozen scope

02

Group, entity and area are pinned: not even the agent itself can ask for a different context.

Credential issued to the identity

03

The credential is issued in the name of an identity and inherits its scope: change the profile's area and what the agent sees changes, without touching the integration.

Scales with your portfolio

04

Twenty contexts are twenty isolated sets of profiles, governed from a single place. You define the profile once and assign it to any agent, today or next year.

In a regulated organization, it guarantees an agent never leaks anyone a document they weren't supposed to see.

How to use it

MCP, REST, CLI or SDK. Use it from wherever you work.

Four surfaces, on every plan. The message is one of neutrality; the path we show first is MCP.

All six interfaces —web chat, MCP, REST, Python SDK, TypeScript SDK and CLI— are on every plan, the free one included. RAGfly Desktop is separate, from Starter up.

MCP connectionNo integration project.
{
  "mcpServers": {
    "ragfly": {
      "url": "https://api.ragfly.ai/mcp/sse",
      "headers": { "Authorization": "Bearer rf_xxxxxxxxxx" }
    }
  }
}
NOT JUST CONTEXT

Not just context: action

The agent doesn't stop at retrieving chunks: it invokes operations over sets of documents by prompt, from your own agent.

«Download these 500 documents and put them where I tell you.»

RAGfly is a toolset of document operations that the agent calls, not just a context provider.

NO LOCK-IN

Your models, your database, your documents.

You're always in control.

MODELOWNER · YOU

Your models

With parsing-BYOK and embedding-BYOK, processing runs on your credentials, against Anthropic, OpenAI or Google, under your contract and your invoice.

DATABASEOWNER · YOU

Your database

Vectors can live in your own database instead of ours: on Scale, your own Supabase —the same engine that runs RAGfly, Postgres with pgvector—; on Enterprise, other engines.

DOCUMENTSOWNER · YOU

Your documents

You never lose control of your documents: we never store them.

Your models never see the document and your vectors live on your infrastructure.

Plans

Clear plans for simple, private and precise RAG.

PRICES & QUOTAS · USD/MONTH · LIVE SOURCE

Free

Plan gratuito permanente para conocer el valor completo de RAGfly con límites mensuales chicos.

Free
  • 1,000 Active corpus/month
  • 1,500 Pure retrievals/month
  • 30 Verified Answers/month
  • 5 Agentic Retrieval/month
  • Other limits: Entities: 2 · Areas: 4 · Workspaces: 5 · Synthesis (source pages): 1,500 · ASISTENCIAS: Custom
Start free

Starter

Primer agente documental en produccion para un equipo pequeno.

$39/month
  • 10,000 Active corpus/month
  • 5,000 Pure retrievals/month
  • 300 Verified Answers/month
  • 50 Agentic Retrieval/month
  • Other limits: Entities: 5 · Areas: 12 · Workspaces: 10 · Synthesis (source pages): 10,000 · ASISTENCIAS: Custom
  • Add-ons: corpus +10,000 for $10/month · Augmented RAG +500 for $35 · Agentic RAG +50 for $59.
Get started
Recommended

Growth

Plan para operar varios clientes, areas o flujos con control de visibilidad.

$149/month
  • 50,000 Active corpus/month
  • 25,000 Pure retrievals/month
  • 1,500 Verified Answers/month
  • 150 Agentic Retrieval/month
  • Other limits: Entities: 25 · Areas: 50 · Workspaces: 50 · Synthesis (source pages): 50,000 · ASISTENCIAS: Custom
  • Add-ons: corpus +10,000 for $10/month · Augmented RAG +500 for $35 · Agentic RAG +50 for $59.
Get started

Scale

Alto volumen y arquitectura extensible con Supabase vectorial propio (BYO).

$590/month
  • 250,000 Active corpus/month
  • 100,000 Pure retrievals/month
  • 5,000 Verified Answers/month
  • 500 Agentic Retrieval/month
  • Other limits: Entities: 150 · Areas: 200 · Workspaces: 200 · Synthesis (source pages): 200,000 · ASISTENCIAS: Custom
  • Add-ons: corpus +10,000 for $10/month · Augmented RAG +500 for $35 · Agentic RAG +50 for $59.
Get started

Enterprise / Sovereign

Plan custom para despliegue gestionado, soberano u on-premise.

Custom
  • Custom Active corpus
  • Custom Pure retrievals
  • Custom Verified Answers
  • Custom Agentic Retrieval
  • Other limits: Entities: Custom · Areas: Custom · Workspaces: Custom · Synthesis (source pages): Custom · ASISTENCIAS: Custom
Talk to us
Free is permanent, needs no card and has monthly allowances. Paid plans support manual packs and recurring corpus capacity; there is no automatic overage. Included processing is subject to reasonable use to protect shared performance.
CLOSING

Point RAGfly at your documents and give any agent the exact context it needs — production-ready and isolated by context.

We build for agents in production, not for demos.

We measure what matters: agents and applications serving retrieval from RAGfly. Not chat users, not documents uploaded.

FREQUENTLY ASKED QUESTIONS

What people ask us before starting

What is RAGfly?

A RAG service for AI agents: it turns any document corpus —thousands or tens of thousands, scans included— into a secure, multi-tenant retrieval base that's ready for production, without building or maintaining a RAG pipeline.

Who is it for?

For developers, consultancies and integrators building AI agents on private documents —sometimes for many different clients— who don't want to build or maintain a RAG pipeline at scale.

How much does it cost?

Plans in USD: Free $0, Starter $39/month, Growth $149/month, Scale $590/month and Enterprise by agreement. Each plan shows active corpus, Simple Retrieval, Augmented RAG, Agentic RAG and entities; paid plans support manual add-ons and have no automatic overage. MCP, REST, CLI and SDK are available on every plan.

How does it work?

You point RAGfly at your documents directory; the system scans, vectorizes and indexes them automatically. Then your agent retrieves by meaning and gets answers with citations to the source via MCP, REST or CLI, isolated by context and with the data handling that matches the deployment mode.

How does an AI agent consume RAGfly?

Through a remote MCP server (SSE or HTTP) or the RAGfly Desktop CLI, authenticating with a JWT or an API Key. The catalog of operations is published at `https://ragfly.ai/agents.json` and `https://ragfly.ai/llms-full.txt`.

What does RAGfly process, and where does my data end up?

RAGfly processes document information on the way in and relevant chunks on the way out. In Cloud it processes the original file but does not retain it after ingestion; chunks, embeddings and metadata persist. With Desktop, the original and its extraction are processed inside your network, although derivatives may be synchronized. On sovereign plans, the database and the derivatives at rest belong to the customer; RAGfly accesses the chunks it needs to retrieve and deliver context.

How is one customer's data isolated from another's?

Multi-tenant by default: each customer lives in a corpus isolated at the database level, with a Groups → Entities → Areas structure. A consultancy or integrator can serve dozens of customers from a single platform without a single record crossing over.

Can I use my own vector database and my own model?

It depends on the plan. Growth enables parsing-BYOK, embedding-BYOK and agentic-BYOK; Scale adds a customer-owned Supabase vector project; Enterprise supports customer LLMs and compatible vector databases.

Which AI models does it work with — Claude, GPT, Gemini?

RAGfly is model-agnostic. It delivers retrieval with citations and permissions; your agent runs on whichever LLM you prefer —Claude, GPT, Gemini or another— and consumes RAGfly via MCP, REST or CLI.

Which formats are supported? Does it work with scanned PDFs (OCR)?

PDF, Word, Excel and plain text, including scanned documents thanks to first-rate OCR and table and layout understanding. The focus is text and complex documents, not audio or video.

Can I use it with regulated data (legal, healthcare, government)?

Yes. Desktop processes the original file and its extraction inside the customer's network. On sovereign plans, the customer owns the database where the derivatives persist. RAGfly still processes chunks on the way out to produce retrieval; a full in-situ runtime is the mode that keeps that processing inside the customer's infrastructure too.

How long does it take to go from a directory to an agent answering with citations?

You point RAGfly at a directory and, with automatic ingestion, vectorization and indexing, you can have an agent answering with citations in a short time, with no human help — and without building or maintaining a RAG pipeline at scale.

Not ready to sign up? Join the early-access list and we'll let you know.

Prefer to write to us? info@ragfly.ai