OpenCrawling ran two code reviews on a workspace and the local model barely cost them nothing
Published
Signal category
Technology & Infrastructure
Quote
“almost local”
— Michael Cizmar|OpenCrawling team
Company
OpenCrawling
Open Source Enterprise Content Crawling & Streaming Engine for RAG, Vector Search, and LLM Agents.
- Industry
- Software Development
- Company size
- 2 employees
OpenCrawling is the next-generation open-source data federation and ingestion engine, built to securely bridge the gap between fragmented enterprise data silos and modern Artificial Intelligence. As Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) architectures become the standard, the true bottleneck is no longer the AI itself, but the speed, quality, and security of the context provided to it. Traditional web crawlers and legacy ingestion pipelines are blocking monoliths—too resource-heavy for real-time AI needs and prone to stripping away critical document security permissions during extraction. OpenCrawling solves this by rethinking data ingestion from the ground up: 🚀 Cloud-Native & High-Throughput: Built natively on Java 25 and Spring Boot 4, the engine abandons legacy thread pools. By leveraging Virtual Threads and structured concurrency, it achieves zero-blocking, high-throughput I/O across thousands of simultaneous connections with a minimal memory footprint. 🔒 Zero-Trust AI & ACL Synchronization: AI should never compromise enterprise security. OpenCrawling natively reads and maps Access Control Lists (ACLs) from source repositories (ECMs, CRMs, Databases), transferring user permissions directly into Vector Databases. We ensure that AI agents only retrieve and generate answers based on documents the specific user is explicitly authorized to see. 🧩 Pluggable & Standardized (OIS): OpenCrawling introduces the Open Ingestion Standard (OIS). Through decoupled Service Provider Interfaces (SPIs) and native Model Context Protocol (MCP) integration, developers can easily build custom source and target connectors. The engine acts as a live, secure MCP Server, providing LLMs with an authorized, real-time context window into the enterprise. We are an open-source community dedicated to establishing the new architectural standard for Enterprise Information Management in the AI era.
Founded 2026