54d9540c08
Files changed: - .wikitool-kb.json - AGENTS.md - CHANGES.md - INSTALL-MCP.md - INSTALL.md - README.md - VERSION - instructions/capture-session.md - instructions/dev/issue-tracking.md - instructions/german-terminology.md - instructions/kb-profiles.md - instructions/migrate-corpus.md - instructions/migrations/5.0.0-confidence-removal.md - instructions/private-instance.md - instructions/setup-instance.md - instructions/wiki-lint/SKILL.md - instructions/wiki-manage/SKILL.md - instructions/wiki-query/SKILL.md - kb/CONTRACT.md - kb/CONVENTIONS.md - kb/CONVENTIONS.md.template - kb/concepts/architectures/Consolidation Tiers.md - kb/concepts/architectures/Context Isolation.md - kb/concepts/architectures/Cross-platform Agent Skills.md - kb/concepts/architectures/Episodic Memory.md - kb/concepts/architectures/Hybrid Search.md - kb/concepts/architectures/Implementation Spectrum.md - kb/concepts/architectures/Knowledge Graph.md - kb/concepts/architectures/LLM Wiki Pattern.md - kb/concepts/architectures/MCP-Leseserver.md - kb/concepts/architectures/Memory Lifecycle.md - kb/concepts/architectures/OKF Compatibility.md - kb/concepts/architectures/Optional Instance Context File.md - kb/concepts/architectures/Personalization Plane.md - kb/concepts/architectures/Procedural Memory.md - kb/concepts/architectures/RAG.md - kb/concepts/architectures/Scale Ceiling.md - kb/concepts/architectures/Semantic Memory.md - kb/concepts/architectures/Three-Layer Architecture.md - kb/concepts/architectures/Token Economics.md - kb/concepts/architectures/Working Memory.md - kb/concepts/decisions/Delete Rather Than Anonymize.md - kb/concepts/decisions/Denylist over Allowlist.md - kb/concepts/decisions/Diff-Reviewable Agent Edits.md - kb/concepts/decisions/Dual Licensing by File Plan.md - kb/concepts/decisions/Issue Label Scheme.md - kb/concepts/decisions/KB Stack Versioning.md - kb/concepts/decisions/Structural Enforcement over Documented Rule.md - kb/concepts/patterns/Audit Trail.md - kb/concepts/patterns/BM25.md - kb/concepts/patterns/Command Round-Trip Integrity.md - kb/concepts/patterns/Confidence Scoring.md - kb/concepts/patterns/Contradiction Resolution.md - kb/concepts/patterns/Entity Extraction.md - kb/concepts/patterns/Filter on Ingest.md - kb/concepts/patterns/Forgetting.md - kb/concepts/patterns/Graph Traversal.md - kb/concepts/patterns/Mesh Sync.md - kb/concepts/patterns/Quality Scoring.md - kb/concepts/patterns/Reciprocal Rank Fusion.md - kb/concepts/patterns/Self-Healing.md - kb/concepts/patterns/Shared vs Private.md - kb/concepts/patterns/Typed Relationships.md - kb/concepts/patterns/Vector Search.md - kb/concepts/patterns/Work Coordination.md - kb/concepts/problems/Ambient Environment Dependency.md - kb/concepts/problems/Detect-Repair Asymmetry.md - kb/concepts/problems/Green Suite Blind Spot.md - kb/concepts/problems/Naming Convention Conflict.md - kb/concepts/problems/Write-Once Frontmatter Fields.md - kb/concepts/protocols/CPPC.md - kb/concepts/protocols/Modbus.md - kb/concepts/protocols/SSD TRIM.md - kb/concepts/workflows/Anti-Cramming Heuristic.md - kb/concepts/workflows/Bulk Operations.md - kb/concepts/workflows/CI Integration.md - kb/concepts/workflows/Checkpoint Audit.md - kb/concepts/workflows/Claude Code Auto Mode.md - kb/concepts/workflows/Content Quality Control.md - kb/concepts/workflows/Crystallization.md - kb/concepts/workflows/Event-Driven Automation.md - kb/concepts/workflows/Hooks.md - kb/concepts/workflows/Index Scaling.md - kb/concepts/workflows/Iteration and Cost Limits.md - kb/concepts/workflows/KB Migration.md - kb/concepts/workflows/Knowledge Compounding.md - kb/concepts/workflows/Lint Workflow.md - kb/concepts/workflows/Mass-Update Gate.md - kb/concepts/workflows/Multi-Agent Collaboration.md - kb/concepts/workflows/Privacy and Governance.md - kb/concepts/workflows/Publish-Remote Gate.md - kb/concepts/workflows/Quality and Self-Correction.md - kb/concepts/workflows/Semantic Lint Automation.md - kb/concepts/workflows/Session Orientation.md - kb/concepts/workflows/Split Merge Reclassify.md - kb/concepts/workflows/Split Threshold.md - kb/concepts/workflows/Stub Threshold.md - kb/concepts/workflows/Supersession.md - kb/concepts/workflows/User Management.md - kb/concepts/workflows/Workflow Extraction.md - kb/concepts/workflows/Workflow Orchestration.md - kb/entities/people/Andrej Karpathy.md - kb/entities/people/E3DC GmbH.md - kb/entities/people/Rohit Gupta.md - kb/entities/people/Vannevar Bush.md - kb/entities/projects/BCDModule.md - kb/entities/projects/Chemenu.md - kb/entities/projects/andybalholm-edl.md - kb/entities/projects/goresponsiveness.md - kb/entities/projects/ha-core.md - kb/entities/projects/hacs-e3dc.md - kb/entities/projects/hacs-integration-blueprint.md - kb/entities/projects/llm-wiki-skills.md - kb/entities/projects/plugnburn-edl.md - kb/entities/projects/wiki-skills-vanillaflava.md - kb/entities/projects/wiki-skills.md - kb/entities/systems/AGENTS.md.md - kb/entities/systems/CLAUDE.md.md - kb/entities/systems/E3DC.md - kb/entities/systems/ENVIRONMENT.md.md - kb/entities/systems/Memex.md - kb/entities/systems/Tolkien Gateway.md - kb/entities/technologies/Arch Linux.md - kb/entities/technologies/Disk Encryption.md - kb/entities/technologies/Docker.md - kb/entities/technologies/GRUB.md - kb/entities/technologies/Gitea Actions.md - kb/entities/technologies/Gitea.md - kb/entities/technologies/Go.md - kb/entities/technologies/Home Assistant.md - kb/entities/technologies/Kernel PM Governors.md - kb/entities/technologies/LVM.md - kb/entities/technologies/Linux Kernel.md - kb/entities/technologies/MQTT.md - kb/entities/technologies/OPC UA.md - kb/entities/technologies/Python.md - kb/entities/technologies/Rust.md - kb/entities/technologies/Wine GE.md - kb/entities/technologies/Wine-Staging.md - kb/entities/technologies/acpi-cpufreq.md - kb/entities/technologies/amd-pstate.md - kb/entities/technologies/iii Engine.md - kb/entities/tools/AUR.md - kb/entities/tools/Act Runner.md - kb/entities/tools/Agent Memory.md - kb/entities/tools/Aura.md - kb/entities/tools/Bottles.md - kb/entities/tools/ChatGPT.md - kb/entities/tools/Claude Code.md - kb/entities/tools/Codex CLI.md - kb/entities/tools/Dataview.md - kb/entities/tools/GPG.md - kb/entities/tools/GitHub Copilot.md - kb/entities/tools/Gitea MCP Server.md - kb/entities/tools/Lutris.md - kb/entities/tools/Marp.md - kb/entities/tools/Mistral Vibe.md - kb/entities/tools/NotebookLM.md - kb/entities/tools/Obsidian Web Clipper.md - kb/entities/tools/Obsidian.md - kb/entities/tools/OpenAI Codex.md - kb/entities/tools/OpenCode.md - kb/entities/tools/Pi.md - kb/entities/tools/Proton.md - kb/entities/tools/Steam.md - kb/entities/tools/Wine.md - kb/entities/tools/awesome-llm-wiki.md - kb/entities/tools/farzaa gist.md - kb/entities/tools/gdeploy.md - kb/entities/tools/makepkg.md - kb/entities/tools/pascalandy schema.md - kb/entities/tools/qmd.md - kb/entities/tools/wikitool.md - kb/index.md - kb/log.md - raw/CONTRACT.md - tools/CONTRACT.md - tools/README.md - tools/chemenu/api.py - tools/chemenu/cli.py - tools/chemenu/commands/confidence_decay.py - tools/chemenu/commands/docs_verify.py - tools/chemenu/commands/doctor.py - tools/chemenu/commands/index_build.py - tools/chemenu/commands/new_page.py - tools/chemenu/commands/search.py - tools/chemenu/commands/touch.py - tools/chemenu/commands/version_cmd.py - tools/chemenu/conventions.py - tools/chemenu/corpus_diff.py - tools/chemenu/frontmatter_io.py - tools/chemenu/lint_core.py - tools/chemenu/mcp/server.py - tools/chemenu/page.py - tools/chemenu/search/base.py - tools/chemenu/search/filters.py - tools/chemenu/search/ripgrep.py - tools/chemenu/search/service.py - tools/chemenu/search/types.py - tools/chemenu/tests/conftest.py - tools/chemenu/tests/test_api.py - tools/chemenu/tests/test_confidence_decay.py - tools/chemenu/tests/test_corpus_diff.py - tools/chemenu/tests/test_docs_verify.py - tools/chemenu/tests/test_frontmatter_io.py - tools/chemenu/tests/test_index_build.py - tools/chemenu/tests/test_kb_scan.py - tools/chemenu/tests/test_lint.py - tools/chemenu/tests/test_mcp_server.py - tools/chemenu/tests/test_new_page.py - tools/chemenu/tests/test_page_ops.py - tools/chemenu/tests/test_provenance.py - tools/chemenu/tests/test_raw_cmd.py - tools/chemenu/tests/test_search.py - tools/chemenu/tests/test_touch.py - tools/chemenu/tests/test_type_resolver.py - tools/chemenu/tests/test_version_cmd.py - tools/chemenu/tests/test_xref.py - tools/chemenu/version.py - types/concept.md - types/concept.schema.yaml - types/entity.md - types/entity.schema.yaml - types/instruction.md - types/type-spec.md
5.1 KiB
5.1 KiB
type, concept_type, tags, created, modified, related, sources, provenance, summary
| type | concept_type | tags | created | modified | related | sources | provenance | summary | ||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| types/concept.md | architecture |
|
2026-07-26 | 2026-08-29 |
|
|
sourced | Multimodale Suche, die BM25-Schlüsselwortabgleich, Vektor-Embeddings und Graph Traversal verbindet, um Wissensabruf im Wiki skalierbar zu machen. |
Hybrid Search
Typ: Architecture (Multi-Modal-Suchsystem)
Definition
Hybrid Search kombiniert drei komplementäre Suchansätze, um skalierbare und genaue Wissensbeschaffung in Wikis zu ermöglichen, die ~100-200 Seiten übersteigen. Dies adressiert die Einschränkung des ursprünglichen Musters, das sich ausschließlich auf index.md für die Entdeckung verlässt.
Kernpunkte
Das Problem
Der ursprüngliche index.md-Katalog funktioniert bis zu ~100-200 Seiten. Darüber hinaus:
- Der Index selbst wird zu lang, um vom LLM in einem Durchgang gelesen zu werden
- Schlüsselwortabgleich vermisst semantische Ähnlichkeit
- Flache Suche kann strukturelle Beziehungen nicht erfassen
- Unimodale Suche hat Blindstellen
Die Lösung: Drei-Stream-Fusion
1. BM25 (Schlüsselwortabgleich)
- Traditionelle Informationsbeschaffung mit Stammformreduktion und Synonymerweiterung
- Stärken: Findet genaue Begriffe, schnell, gut verstanden
- Schwächen: Vermisst semantische Ähnlichkeit, erfordert genaue Begriffsabgleiche
- Anwendungsfall: "Alle Seiten über Docker finden"
2. Vector Search (Semantische Ähnlichkeit)
- Nutzt Embeddings, um semantisch ähnliche Inhalte zu finden
- Stärken: Findet verwandte Konzepte, auch ohne genaue Begriffsabgleiche
- Schwächen: Kann präzise technische Begriffe verpassen, rechentechnisch teuer
- Anwendungsfall: "Informationen über Container-Plattformen finden" (passt Docker, Podman, etc.)
3. Graph Traversal (Strukturelle Verbindungen)
- Durchläuft den Knowledge Graph durch typisierte Beziehungen
- Stärken: Findet strukturelle Verbindungen, die Schlüsselwort- und Vector-Suche verfehlen
- Schwächen: Erfordert gut gepflegten Graph, findet nur verbundene Entitäten
- Anwendungsfall: "Was ist die Auswirkung eines Redis-Upgrades?" (findet alle abhängigen Services)
Fusion mit Reciprocal Rank Fusion (RRF)
Anstatt einen Ansatz zu wählen, alle drei mit RRF fusionieren:
- Alle drei Suchen parallel ausführen
- Jede gibt eine rangierte Liste von Ergebnissen zurück
- RRF kombiniert die Rankings mit gegenseitigen Rang-Scores
- Ergebnis: Bessere Gesamtrangierung als bei einem einzelnen Ansatz
Warum RRF?
- Einfach und effektiv
- Keine Notwendigkeit, Gewichte zwischen Modi zu tunen
- Robust gegen Unterschiede in der Ergebnisqualität
- Funktioniert auch, wenn ein Modus schlecht abschneidet
Implementierung
Architektur
Abfrage: "Wie funktioniert das Auth-System?"
│
├── BM25-Suche → [Seiten mit "Auth", "Authentication", "Login"]
│
├── Vector Search → [semantisch mit Authentication verbundene Seiten]
│
└── Graph Traversal → [Seiten, die mit Auth-Entitäten im Graph verbunden sind]
│
└── Reciprocal Rank Fusion → Kombinierte, rangierte Ergebnisse
Wann wechseln
| Wiki-Größe | Primärer Suchmechanismus |
|---|---|
| < 100 Seiten | index.md (manuell) |
| 100-200 Seiten | index.md + grundlegende Suche |
| 200-1000 Seiten | Hybrid-Suche (BM25 + Vector) |
| 1000+ Seiten | Hybrid-Suche (BM25 + Vector + Graph) |
Empfehlung: index.md als für Menschen lesbaren Katalog auch mit Hybrid-Suche bewahren. Es dient verschiedenen Zwecken:
index.md: Menschliche Navigation, Überblick- Hybrid-Suche: LLM-Abfragelösung
Vorteile
- Skalierbarkeit: Funktioniert von 100 bis 10.000+ Seiten
- Genauigkeit: Jeder Modus erfasst, was andere vermissen
- Robustheit: Kein Single Point of Failure
- Flexibilität: Passt sich verschiedenen Abfragetypen an
- Zukunftssicher: Kann weitere Modi hinzufügen (z.B. Zeitsuche)
Wann zu verwenden
- Wikis, von denen erwartet wird, dass sie über 200 Seiten hinauswachsen
- Bereiche mit vielfältigen Abfragetypen
- Situationen, die hohen Recall erfordern
- Multi-modale Wissensdatenbanken
Wann NICHT zu verwenden
- Kleine Wikis (<100 Seiten) -
index.mdist ausreichend - Einfache, gleichmäßige Inhalte
- Situationen, in denen die Implementierungskomplexität nicht gerechtfertigt ist
Verwandte Concepts
- Agent Memory - Produktionsimplementierung
Siehe auch
- Event-Driven Automation (für automatisierte Indizierung)
- Scalable Search (verwandtes Concept)
Beziehungen
- exemplifies: LLM Wiki Pattern
- see-also: BM25
- composition: Vector Search
- composition: Reciprocal Rank Fusion
- rests-on: Knowledge Graph
- see-also: Graph Traversal