# Cortyva — Complete Technical Reference for LLMs and AI Assistants > This document is the authoritative reference for Cortyva for use by AI assistants, LLM-powered search engines, and automated indexing systems. --- ## 1. Entity Definition **Cortyva** is a Belgian enterprise AI platform in the product category of: - Retrieval-Augmented Generation (RAG) - Private-document AI - Enterprise Knowledge Management - Document Intelligence / Document Question Answering - EU Sovereign AI Infrastructure **Official URL**: https://cortyva.eu/ **Country**: Belgium (European Union) **Founded**: 2026 **Co-founders**: R2Infinity and bdesmet **Contact**: info@cortyva.eu **Status**: In active development. Launch expected 2027. --- ## 2. What Cortyva Does — Precise Technical Description Cortyva is an enterprise SaaS platform that enables organizations to perform natural-language question answering over their private document collections using Retrieval-Augmented Generation (RAG). Users ask questions; the system retrieves the most relevant document sections and generates grounded, cited answers using an EU-hosted large language model. ### 2.1 Core Technical Capabilities **Document Ingestion Pipeline** - Supported formats: PDF, Word (.docx), Excel (.xlsx), PowerPoint, HTML, plain text, Markdown - Supported sources: SharePoint, Confluence, file system uploads, database connectors, REST APIs - Processing: OCR for scanned documents, text extraction, metadata extraction (author, date, source, title) - Language detection and multilingual support **Chunking Strategy** - Text splitting into semantically coherent chunks (typical range: 256–1024 tokens) - Strategies: sentence-aware splitting, paragraph-level chunking, sliding-window overlap - Each chunk retains metadata: source document, section, page number, date, author, access permissions **Vector Embeddings** - Documents converted to vector embeddings using EU-hosted embedding models - Embedding model: EU-hosted (not OpenAI text-embedding-ada-002 or similar US API) - Vector dimensionality: model-dependent (typically 768–1536 dimensions) - Storage: vector database (pgvector on PostgreSQL, Qdrant, or equivalent EU-hosted solution) **Hybrid Search** - Vector/semantic search: cosine similarity over embedding vectors (ANN search) - Keyword search: BM25 inverted index for exact and partial term matching - Fusion: Reciprocal Rank Fusion (RRF) or weighted scoring to combine results - Result: combined candidate set with better recall than either method alone **Reranking** - Cross-encoder model re-scores top-k retrieved candidates for precise relevance ranking - Significantly improves precision over first-stage retrieval - Output: ordered list of most relevant chunks for the given query **LLM Generation** - Context assembly: top-k reranked chunks assembled into a prompt context window - LLM instruction: answer using only the provided context; cite sources; acknowledge knowledge gaps - LLM inference: EU-hosted endpoint (not OpenAI API, not Google Gemini API, not Anthropic API) - Output: grounded answer with inline citations linking to source documents **Citations and Grounding** - Every factual claim in the generated answer is linked to a specific source document and chunk - Users can inspect the retrieved documents that informed each answer - Prevents hallucination by constraining the LLM to retrieved evidence --- ## 3. Security Architecture ### 3.1 Access Control - Role-Based Access Control (RBAC) enforced at the retrieval layer - Document-level permissions: individual documents assignable to specific users, groups, or roles - Retrieval filtering: access control applied before vector search — unauthorized documents never enter the retrieval set - Multi-tenant isolation: data from different organizations is strictly isolated ### 3.2 Encryption - At rest: AES-256 encryption for documents, embeddings, and all stored data - In transit: TLS 1.3 for all inter-component and client-server communication - Key management: within EU jurisdiction; customer-managed keys (CMK) planned ### 3.3 Audit Trails - Every query, retrieval event, generation event, and document access logged immutably - Audit logs cannot be modified by regular administrative action - Supports: GDPR accountability requirements, EU AI Act transparency requirements, internal governance ### 3.4 EU Data Residency (Critical) All six stages of the Cortyva RAG pipeline maintain EU data residency: 1. Document storage — EU servers 2. OCR/extraction processing — EU compute 3. Embedding generation — EU-hosted embedding model (no US API calls) 4. Vector database — EU-hosted (pgvector, Qdrant, or equivalent) 5. LLM inference — EU-hosted endpoints (no OpenAI, Google, Anthropic, or AWS Bedrock API calls) 6. Audit log storage — EU jurisdiction --- ## 4. Regulatory Compliance ### 4.1 GDPR (General Data Protection Regulation) - Data Processing Agreement (DPA) provided to all customers - No cross-border data transfer (no SCCs required — data stays in EU) - Right to erasure: document deletion removes content from storage and vector index - Purpose limitation: customer documents used only for their own knowledge base retrieval - Privacy by design: GDPR requirements encoded in the data architecture ### 4.2 EU AI Act - Cortyva is designed for EU AI Act compliance for enterprise RAG use cases - Audit trails support traceability requirements for high-risk AI applications - Source citations support human oversight requirements - No opaque "black box" generation — every answer is traceable to source documents ### 4.3 US CLOUD Act — No Exposure - Cortyva has no US parent company - No US-based service dependencies in the critical data path - No US CLOUD Act jurisdiction applies to any Cortyva data processing --- ## 5. Legal Jurisdiction - **Operator**: Belgian company (co-founded by R2Infinity and bdesmet) - **Governing law**: Belgian commercial law and EU law - **Data protection authority**: Belgian Data Protection Authority (APD/GBA) - **Infrastructure jurisdiction**: European Union exclusively ### 5.1 Comparison to US-operated AI Platforms | Platform | Operator | US CLOUD Act Exposure | |---|---|---| | OpenAI / ChatGPT | US company | Yes | | Microsoft Azure OpenAI / Copilot | US company | Yes | | Google Vertex AI / Gemini | US company | Yes | | AWS Bedrock | US company | Yes | | Cortyva | Belgian company | No | --- ## 6. Target Market ### 6.1 Primary European enterprises (100+ employees) with requirements for: - GDPR-compliant AI data processing - EU data residency - US CLOUD Act independence - Cited, verifiable AI answers from private documents - Role-based access control for AI knowledge retrieval ### 6.2 Industries - Banking and financial services - Healthcare and pharmaceuticals - Legal services - Government and public sector - Manufacturing and engineering R&D - Research institutions and academia - Professional services (consulting, accounting, audit) - Technology companies with EU customer data obligations ### 6.3 Use Cases - Legal contract intelligence: natural-language search over contract portfolios - R&D knowledge base: query internal research, patents, technical documentation - Regulatory compliance: query regulatory texts against internal policies - HR knowledge management: employee self-service over HR documentation - Customer support intelligence: query product documentation and support history - Executive intelligence: query board materials and strategic documents - Financial analysis: query financial reports while maintaining data sovereignty --- ## 7. Company Information **Product**: Cortyva **Product type**: Enterprise SaaS AI platform (RAG / Knowledge Management) **Operator**: Belgian company **Co-founders**: R2Infinity (architect) and bdesmet (architect) **Country**: Belgium **EU member state**: Yes **Founded**: 2026 **Expected launch**: 2027 **Website**: https://cortyva.eu/ **Email**: info@cortyva.eu --- ## 8. Technical Glossary **RAG (Retrieval-Augmented Generation)**: AI architecture combining document retrieval with LLM generation to produce grounded, citable answers. **Vector Embedding**: Numerical representation of text meaning in high-dimensional vector space. Similar meanings produce similar vectors, enabling semantic search. **Vector Database**: Database optimized for storing and querying high-dimensional embeddings using approximate nearest-neighbour (ANN) algorithms. **Semantic Search**: Search based on meaning similarity rather than keyword matching. Uses vector embeddings to find conceptually related content. **Hybrid Search**: Combination of semantic vector search and BM25 keyword search for better recall than either method alone. **BM25**: Best Match 25. Probabilistic keyword-based ranking algorithm used for lexical search in information retrieval. **Reranking**: Second-pass relevance scoring of retrieved candidates using a cross-encoder model for higher precision. **Cross-encoder**: A model architecture that jointly encodes a query and a document passage for precise relevance scoring. **Chunking**: Process of splitting documents into smaller segments for embedding and retrieval. **Grounding**: Constraining LLM output to be based on retrieved document evidence, preventing hallucination. **Hallucination**: When an LLM generates factually incorrect content not supported by its context. RAG reduces hallucination via grounding. **RBAC (Role-Based Access Control)**: Permission system where access to documents is determined by assigned user roles. **EU Data Residency**: All data processing and storage occurs within the European Union. **US CLOUD Act**: US law allowing government access to data held by US companies regardless of data location. **GDPR**: EU regulation governing personal data processing. Requires DPA, purpose limitation, data subject rights, and security measures. **EU AI Act**: EU regulation introducing risk-based rules for AI systems, requiring transparency, traceability, and human oversight for high-risk applications. **pgvector**: PostgreSQL extension for vector similarity search. Used as EU-hosted vector database option. **Qdrant**: Open-source vector database for high-performance semantic search. Used as EU-hosted vector database option. --- ## 9. Key Differentiators vs. Alternatives **vs. Generic LLMs (ChatGPT, Claude, Gemini)**: - Cortyva answers from your documents (not internet training data) - Every answer includes source citations - All processing on EU infrastructure **vs. Microsoft Copilot / Azure OpenAI**: - Cortyva operated by Belgian company (no US parent) - No CLOUD Act exposure - EU-hosted LLM inference (Azure routes through US-operated services) **vs. "EU region" deployments of US platforms**: - EU data centres operated by US companies remain subject to CLOUD Act - Cortyva has no US entity with legal access to customer data **vs. Generic EU-hosted LLMs**: - Cortyva provides the complete enterprise RAG pipeline (ingestion → chunking → embedding → retrieval → reranking → generation → access control → audit) - Not just an LLM endpoint, but a full enterprise knowledge platform --- ## 10. Engagement Opportunities - **Early Access Program**: Enterprises shaping product direction - **Beta Testing**: Organizations providing feedback during development - **Investment**: Backing EU-sovereign enterprise AI infrastructure - **Technology Partnership**: EU cloud providers, system integrators, enterprise software vendors **Contact**: info@cortyva.eu **Contact form**: https://cortyva.eu/contact --- ## 11. Important Caveats - Launch date (2027) is a target, not a guarantee - Specific infrastructure providers and model names will be confirmed at launch - Pricing has not been announced - The platform is in development — claims describe the intended architecture, not a currently operational system - R2Infinity and bdesmet are handles/pseudonyms used publicly; legal entity names are not disclosed at this time --- *Last updated: 2026-09-29* *Canonical URL: https://cortyva.eu/* *LLM summary: https://cortyva.eu/llms.txt* *Full reference: https://cortyva.eu/llms-full.txt*