Company intelligence
OpenCrawling
Open Source Enterprise Content Crawling & Streaming Engine for RAG, Vector Search, and LLM Agents.
About OpenCrawling
OpenCrawling is the next-generation open-source data federation and ingestion engine, built to securely bridge the gap between fragmented enterprise data silos and modern Artificial Intelligence. As Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) architectures become the standard, the true bottleneck is no longer the AI itself, but the speed, quality, and security of the context provided to it. Traditional web crawlers and legacy ingestion pipelines are blocking monoliths—too resource-heavy for real-time AI needs and prone to stripping away critical document security permissions during extraction. OpenCrawling solves this by rethinking data ingestion from the ground up: 🚀 Cloud-Native & High-Throughput: Built natively on Java 25 and Spring Boot 4, the engine abandons legacy thread pools. By leveraging Virtual Threads and structured concurrency, it achieves zero-blocking, high-throughput I/O across thousands of simultaneous connections with a minimal memory footprint. 🔒 Zero-Trust AI & ACL Synchronization: AI should never compromise enterprise security. OpenCrawling natively reads and maps Access Control Lists (ACLs) from source repositories (ECMs, CRMs, Databases), transferring user permissions directly into Vector Databases. We ensure that AI agents only retrieve and generate answers based on documents the specific user is explicitly authorized to see. 🧩 Pluggable & Standardized (OIS): OpenCrawling introduces the Open Ingestion Standard (OIS). Through decoupled Service Provider Interfaces (SPIs) and native Model Context Protocol (MCP) integration, developers can easily build custom source and target connectors. The engine acts as a live, secure MCP Server, providing LLMs with an authorized, real-time context window into the enterprise. We are an open-source community dedicated to establishing the new architectural standard for Enterprise Information Management in the AI era.
Verified activity
Signals from OpenCrawling
2 published signals
Products & Services
OpenCrawling released the Apache Solr Output Connector, which natively streams enterprise document repositories into Apache Solr 10.x and 9.x.
Reported by Michael Cizmar
Technology & Infrastructure
OpenCrawling ran two code reviews on a workspace and the local model barely cost them nothing
Reported by Michael Cizmar