Collection
Research and write from your own sources
Deep-research pipelines, drafting tools and retrieval over your own documents.
29 tools · showing 1–24
LlamaIndex
MemoryData Wrangling
run-llama
Open-source data framework for building LLM and agentic apps over your own data, with 300+ integrations
ragvector-dbocr
ToolRAGFlow
Memory
infiniflow
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
ragdocument-processingknowledge-basecontext-engine
ToolDocling
Data Wrangling
docling-project
Parses PDFs and many other document formats into structured Markdown or JSON, with layout, table, formula, and OCR understanding
productivitydocument-processingocrragmarkdown
OpenRAG
CodingMemory
langflow-ai
OpenRAG is a comprehensive, single package Retrieval-Augmented Generation platform built on Langflow, Docling, and Opensearch.
ragdocument-processingsemantic-searchlangflow
ToolOpenDataLoader PDF
Data Wrangling
opendataloader-project
Extract Markdown, JSON (with bounding boxes), and HTML from any PDF.
document-processingocraccessibilityrag
ToolPaddleOCR
Data Wrangling
PaddlePaddle
OCR toolkit that converts PDFs and images into structured Markdown or JSON, with text, table, formula, and layout recognition
productivityocrdocument-processingtable-recognitionvision-language-model
ToolPageIndex
MemoryResearch
VectifyAI
PageIndex: Document Index for Vectorless, Reasoning-based RAG.
ragvectorless-retrievaldocument-processingmcp-server
PlatformSuperlinked Inference Engine (SIE)
InferenceDeploymentData Wrangling
superlinked
Self-hosted Kubernetes inference cluster serving LLMs, embeddings, rerankers, OCR and vision models with cluster-wide batching.
inference-serverkubernetesrerankingocr
Toolai.adeu/adeu
Data Wrangling
dealfluence
docxtrack-changesdocument-processingredline
PlatformBisheng
Monitoring
dataelement
BISHENG is an open LLM devops platform for next generation Enterprise AI applications. Powerful and comprehensive features include: GenAI wo
ragorchestrationfine-tunedocument-processing
ToolOfficeCLI
Data Wrangling
iOfficeAI
OfficeCLI is the first and best Office suite purpose-built for AI agents to read, edit, and automate Word, Excel, and PowerPoint files. Fre
productivitydocxdocument-processingpowerpointcli
ToolSkill_Seekers
Data Wrangling
yusufkaraaslan
The data layer for AI systems. Turns documentation sites, GitHub repos, PDFs, videos, notebooks, wikis, into structured knowledge assets, ready to power AI Skills, RAG pipelines, and AI coding assistants.
ragclaudegeminidocument-processing
ToolMinerU
Data Wrangling
opendatalab
Converts PDFs and Office documents into LLM-ready Markdown and JSON with layout analysis, OCR and structure extraction
productivitydocument-processingocr
ToolZotero MCP
MemoryData Wrangling
cookjohn
It's a plugin extension in Zotero. Zotero MCP Plugin enables integration between AI assistants and Zotero through MCP. Zotero MCP Plugin 是一
zoterocitation-managementliterature-reviewsemantic-search
ToolOpenKB
Data WranglingMemory
VectifyAI
VectifyAI’s OpenKB compiles raw documents into an interlinked Markdown knowledge base before agents query it. PDFs, docs
knowledge-baseragdocument-processingcli
ToolPullMD
Data WranglingResearch
AeternaLabsHQ
Self-hosted URL- and file-to-Markdown service for humans and AI agents - web pages, documents, images, audio, YouTube. PWA + REST + MCP + Cl
html-to-markdownocryoutube-transcriptdocument-conversion
ToolContextGem
Data Wrangling
shcherbak-ai
Ask an LLM which clause a finding came from and it writes you a citation that reads right and points nowhere. ContextGem
document-processingstructured-outputcontract-analysis
ToolSurya
Data Wrangling
datalab-to
OCR, layout analysis, reading order and table recognition in 90+ languages.
researchdocument-processingocr
ToolLLM Graph Builder
MemoryData Wrangling
neo4j-labs
Turns unstructured documents, media, and web pages into a Neo4j knowledge graph using LLMs, with graph-aware chat over the result
knowledge-graphneo4jraglangchain
ToolDoc7
Data Wrangling
magicrew
Turn documents into AI-ready Markdown with visual understanding
document-processingocrvision-language-modelmarkdown
ToolEasy Dataset
Coding
ConardLi
Build LLM fine-tuning and evaluation datasets from documents with parsing, chunking, QA generation, and multi-format export
dataset-generationfine-tuneragevaluation
ToolHiring Agent
Data Wrangling
interviewstreet
Pipeline that scores resumes by extracting structured data from PDFs and enriching it with GitHub signals
recruitingproductivityresume-screeningdocument-processinggithub-signalsollama
ToolHealthy Diet AI Agent
Memory
archie0732
Bun and TypeScript backend for nutrition chat, food-image analysis, and RAG document retrieval
raglangchainknowledge-graphnutrition
ToolDocument Structuring
Data WranglingMemoryCoding
barnetwang
Parses PDF and DOCX files by heading into Markdown chunks in a local SQLite FTS5 store for RAG retrieval
local-firstragsqlitehybrid-search