Files
chemenu/kb/concepts/architectures/Hybrid Search.md
T
torben 7f74303a00
CI / verify (push) Successful in 52s
Release / release (push) Successful in 36s
kb/concepts/ bekommt Areas: layout: fuer concept, Area-Titel aus jedem Type-Spec, Schwellen-Empfehlung im lint (schliesst #59)
Files changed:
- CHANGES.md
- README.md
- VERSION
- kb/concepts/Ambient Environment Dependency.md
- kb/concepts/Anti-Cramming Heuristic.md
- kb/concepts/Audit Trail.md
- kb/concepts/BM25.md
- kb/concepts/Bulk Operations.md
- kb/concepts/CI Integration.md
- kb/concepts/COLLECTION.md
- kb/concepts/CPPC.md
- kb/concepts/Checkpoint Audit.md
- kb/concepts/Claude Code Auto Mode.md
- kb/concepts/Command Round-Trip Integrity.md
- kb/concepts/Confidence Scoring.md
- kb/concepts/Consolidation Tiers.md
- kb/concepts/Content Quality Control.md
- kb/concepts/Context Isolation.md
- kb/concepts/Contradiction Resolution.md
- kb/concepts/Cross-platform Agent Skills.md
- kb/concepts/Crystallization.md
- kb/concepts/Delete Rather Than Anonymize.md
- kb/concepts/Denylist over Allowlist.md
- kb/concepts/Detect-Repair Asymmetry.md
- kb/concepts/Diff-Reviewable Agent Edits.md
- kb/concepts/Dual Licensing by File Plan.md
- kb/concepts/Entity Extraction.md
- kb/concepts/Episodic Memory.md
- kb/concepts/Event-Driven Automation.md
- kb/concepts/Filter on Ingest.md
- kb/concepts/Forgetting.md
- kb/concepts/Graph Traversal.md
- kb/concepts/Green Suite Blind Spot.md
- kb/concepts/Hooks.md
- kb/concepts/Hybrid Search.md
- kb/concepts/INDEX.md
- kb/concepts/Implementation Spectrum.md
- kb/concepts/Index Scaling.md
- kb/concepts/Issue Label Scheme.md
- kb/concepts/Iteration and Cost Limits.md
- kb/concepts/KB Migration.md
- kb/concepts/KB Stack Versioning.md
- kb/concepts/Knowledge Compounding.md
- kb/concepts/Knowledge Graph.md
- kb/concepts/LLM Wiki Pattern.md
- kb/concepts/Lint Workflow.md
- kb/concepts/MCP-Leseserver.md
- kb/concepts/Mass-Update Gate.md
- kb/concepts/Memory Lifecycle.md
- kb/concepts/Mesh Sync.md
- kb/concepts/Modbus.md
- kb/concepts/Multi-Agent Collaboration.md
- kb/concepts/Naming Convention Conflict.md
- kb/concepts/OKF Compatibility.md
- kb/concepts/Optional Instance Context File.md
- kb/concepts/Personalization Plane.md
- kb/concepts/Privacy and Governance.md
- kb/concepts/Procedural Memory.md
- kb/concepts/Publish-Remote Gate.md
- kb/concepts/Quality Scoring.md
- kb/concepts/Quality and Self-Correction.md
- kb/concepts/RAG.md
- kb/concepts/Reciprocal Rank Fusion.md
- kb/concepts/SSD TRIM.md
- kb/concepts/Scale Ceiling.md
- kb/concepts/Self-Healing.md
- kb/concepts/Semantic Lint Automation.md
- kb/concepts/Semantic Memory.md
- kb/concepts/Session Orientation.md
- kb/concepts/Shared vs Private.md
- kb/concepts/Split Merge Reclassify.md
- kb/concepts/Split Threshold.md
- kb/concepts/Structural Enforcement over Documented Rule.md
- kb/concepts/Stub Threshold.md
- kb/concepts/Supersession.md
- kb/concepts/Three-Layer Architecture.md
- kb/concepts/Token Economics.md
- kb/concepts/Typed Relationships.md
- kb/concepts/User Management.md
- kb/concepts/Vector Search.md
- kb/concepts/Work Coordination.md
- kb/concepts/Workflow Extraction.md
- kb/concepts/Workflow Orchestration.md
- kb/concepts/Working Memory.md
- kb/concepts/Write-Once Frontmatter Fields.md
- kb/concepts/architectures/Consolidation Tiers.md
- kb/concepts/architectures/Context Isolation.md
- kb/concepts/architectures/Cross-platform Agent Skills.md
- kb/concepts/architectures/Episodic Memory.md
- kb/concepts/architectures/Hybrid Search.md
- kb/concepts/architectures/Implementation Spectrum.md
- kb/concepts/architectures/Knowledge Graph.md
- kb/concepts/architectures/LLM Wiki Pattern.md
- kb/concepts/architectures/MCP-Leseserver.md
- kb/concepts/architectures/Memory Lifecycle.md
- kb/concepts/architectures/OKF Compatibility.md
- kb/concepts/architectures/Optional Instance Context File.md
- kb/concepts/architectures/Personalization Plane.md
- kb/concepts/architectures/Procedural Memory.md
- kb/concepts/architectures/RAG.md
- kb/concepts/architectures/Scale Ceiling.md
- kb/concepts/architectures/Semantic Memory.md
- kb/concepts/architectures/Three-Layer Architecture.md
- kb/concepts/architectures/Token Economics.md
- kb/concepts/architectures/Working Memory.md
- kb/concepts/decisions/Delete Rather Than Anonymize.md
- kb/concepts/decisions/Denylist over Allowlist.md
- kb/concepts/decisions/Diff-Reviewable Agent Edits.md
- kb/concepts/decisions/Dual Licensing by File Plan.md
- kb/concepts/decisions/Issue Label Scheme.md
- kb/concepts/decisions/KB Stack Versioning.md
- kb/concepts/decisions/Structural Enforcement over Documented Rule.md
- kb/concepts/patterns/Audit Trail.md
- kb/concepts/patterns/BM25.md
- kb/concepts/patterns/Command Round-Trip Integrity.md
- kb/concepts/patterns/Confidence Scoring.md
- kb/concepts/patterns/Contradiction Resolution.md
- kb/concepts/patterns/Entity Extraction.md
- kb/concepts/patterns/Filter on Ingest.md
- kb/concepts/patterns/Forgetting.md
- kb/concepts/patterns/Graph Traversal.md
- kb/concepts/patterns/Mesh Sync.md
- kb/concepts/patterns/Quality Scoring.md
- kb/concepts/patterns/Reciprocal Rank Fusion.md
- kb/concepts/patterns/Self-Healing.md
- kb/concepts/patterns/Shared vs Private.md
- kb/concepts/patterns/Typed Relationships.md
- kb/concepts/patterns/Vector Search.md
- kb/concepts/patterns/Work Coordination.md
- kb/concepts/problems/Ambient Environment Dependency.md
- kb/concepts/problems/Detect-Repair Asymmetry.md
- kb/concepts/problems/Green Suite Blind Spot.md
- kb/concepts/problems/Naming Convention Conflict.md
- kb/concepts/problems/Write-Once Frontmatter Fields.md
- kb/concepts/protocols/CPPC.md
- kb/concepts/protocols/Modbus.md
- kb/concepts/protocols/SSD TRIM.md
- kb/concepts/workflows/Anti-Cramming Heuristic.md
- kb/concepts/workflows/Bulk Operations.md
- kb/concepts/workflows/CI Integration.md
- kb/concepts/workflows/Checkpoint Audit.md
- kb/concepts/workflows/Claude Code Auto Mode.md
- kb/concepts/workflows/Content Quality Control.md
- kb/concepts/workflows/Crystallization.md
- kb/concepts/workflows/Event-Driven Automation.md
- kb/concepts/workflows/Hooks.md
- kb/concepts/workflows/Index Scaling.md
- kb/concepts/workflows/Iteration and Cost Limits.md
- kb/concepts/workflows/KB Migration.md
- kb/concepts/workflows/Knowledge Compounding.md
- kb/concepts/workflows/Lint Workflow.md
- kb/concepts/workflows/Mass-Update Gate.md
- kb/concepts/workflows/Multi-Agent Collaboration.md
- kb/concepts/workflows/Privacy and Governance.md
- kb/concepts/workflows/Publish-Remote Gate.md
- kb/concepts/workflows/Quality and Self-Correction.md
- kb/concepts/workflows/Semantic Lint Automation.md
- kb/concepts/workflows/Session Orientation.md
- kb/concepts/workflows/Split Merge Reclassify.md
- kb/concepts/workflows/Split Threshold.md
- kb/concepts/workflows/Stub Threshold.md
- kb/concepts/workflows/Supersession.md
- kb/concepts/workflows/User Management.md
- kb/concepts/workflows/Workflow Extraction.md
- kb/concepts/workflows/Workflow Orchestration.md
- kb/index.md
- kb/log.md
- tools/CONTRACT.md
- tools/README.md
- tools/chemenu/catalog.py
- tools/chemenu/commands/index_build.py
- tools/chemenu/lint_core.py
- tools/chemenu/tests/conftest.py
- tools/chemenu/tests/test_cite_cmd.py
- tools/chemenu/tests/test_git_publish.py
- tools/chemenu/tests/test_index_build.py
- tools/chemenu/tests/test_lint.py
- tools/chemenu/tests/test_new_page.py
- tools/chemenu/tests/test_provenance.py
- tools/chemenu/tests/test_type_resolver.py
- tools/chemenu/tests/test_xref.py
- types/concept.md
- types/type-spec.md
2026-09-08 10:07:46 +02:00

5.1 KiB

type, concept_type, tags, created, modified, related, sources, confidence, confidence_base, provenance, summary
type concept_type tags created modified related sources confidence confidence_base provenance summary
types/concept.md architecture
search
bm25
vector
graph
scalability
2026-07-26 2026-08-29
exemplifies
LLM Wiki Pattern
see-also
BM25
composition
Vector Search
composition
Reciprocal Rank Fusion
rests-on
Knowledge Graph
see-also
Graph Traversal
Source - LLM Wiki v2
0.90 0.90 sourced Multimodale Suche, die BM25-Schlüsselwortabgleich, Vektor-Embeddings und Graph Traversal verbindet, um Wissensabruf im Wiki skalierbar zu machen.

Hybrid Search

Typ: Architecture (Multi-Modal-Suchsystem)

Definition

Hybrid Search kombiniert drei komplementäre Suchansätze, um skalierbare und genaue Wissensbeschaffung in Wikis zu ermöglichen, die ~100-200 Seiten übersteigen. Dies adressiert die Einschränkung des ursprünglichen Musters, das sich ausschließlich auf index.md für die Entdeckung verlässt.

Kernpunkte

Das Problem

Der ursprüngliche index.md-Katalog funktioniert bis zu ~100-200 Seiten. Darüber hinaus:

  • Der Index selbst wird zu lang, um vom LLM in einem Durchgang gelesen zu werden
  • Schlüsselwortabgleich vermisst semantische Ähnlichkeit
  • Flache Suche kann strukturelle Beziehungen nicht erfassen
  • Unimodale Suche hat Blindstellen

Die Lösung: Drei-Stream-Fusion

1. BM25 (Schlüsselwortabgleich)

  • Traditionelle Informationsbeschaffung mit Stammformreduktion und Synonymerweiterung
  • Stärken: Findet genaue Begriffe, schnell, gut verstanden
  • Schwächen: Vermisst semantische Ähnlichkeit, erfordert genaue Begriffsabgleiche
  • Anwendungsfall: "Alle Seiten über Docker finden"

2. Vector Search (Semantische Ähnlichkeit)

  • Nutzt Embeddings, um semantisch ähnliche Inhalte zu finden
  • Stärken: Findet verwandte Konzepte, auch ohne genaue Begriffsabgleiche
  • Schwächen: Kann präzise technische Begriffe verpassen, rechentechnisch teuer
  • Anwendungsfall: "Informationen über Container-Plattformen finden" (passt Docker, Podman, etc.)

3. Graph Traversal (Strukturelle Verbindungen)

  • Durchläuft den Knowledge Graph durch typisierte Beziehungen
  • Stärken: Findet strukturelle Verbindungen, die Schlüsselwort- und Vector-Suche verfehlen
  • Schwächen: Erfordert gut gepflegten Graph, findet nur verbundene Entitäten
  • Anwendungsfall: "Was ist die Auswirkung eines Redis-Upgrades?" (findet alle abhängigen Services)

Fusion mit Reciprocal Rank Fusion (RRF)

Anstatt einen Ansatz zu wählen, alle drei mit RRF fusionieren:

  1. Alle drei Suchen parallel ausführen
  2. Jede gibt eine rangierte Liste von Ergebnissen zurück
  3. RRF kombiniert die Rankings mit gegenseitigen Rang-Scores
  4. Ergebnis: Bessere Gesamtrangierung als bei einem einzelnen Ansatz

Warum RRF?

  • Einfach und effektiv
  • Keine Notwendigkeit, Gewichte zwischen Modi zu tunen
  • Robust gegen Unterschiede in der Ergebnisqualität
  • Funktioniert auch, wenn ein Modus schlecht abschneidet

Implementierung

Architektur

Abfrage: "Wie funktioniert das Auth-System?"
       │
       ├── BM25-Suche → [Seiten mit "Auth", "Authentication", "Login"]
       │
       ├── Vector Search → [semantisch mit Authentication verbundene Seiten]
       │
       └── Graph Traversal → [Seiten, die mit Auth-Entitäten im Graph verbunden sind]
               │
               └── Reciprocal Rank Fusion → Kombinierte, rangierte Ergebnisse

Wann wechseln

Wiki-Größe Primärer Suchmechanismus
< 100 Seiten index.md (manuell)
100-200 Seiten index.md + grundlegende Suche
200-1000 Seiten Hybrid-Suche (BM25 + Vector)
1000+ Seiten Hybrid-Suche (BM25 + Vector + Graph)

Empfehlung: index.md als für Menschen lesbaren Katalog auch mit Hybrid-Suche bewahren. Es dient verschiedenen Zwecken:

  • index.md: Menschliche Navigation, Überblick
  • Hybrid-Suche: LLM-Abfragelösung

Vorteile

  • Skalierbarkeit: Funktioniert von 100 bis 10.000+ Seiten
  • Genauigkeit: Jeder Modus erfasst, was andere vermissen
  • Robustheit: Kein Single Point of Failure
  • Flexibilität: Passt sich verschiedenen Abfragetypen an
  • Zukunftssicher: Kann weitere Modi hinzufügen (z.B. Zeitsuche)

Wann zu verwenden

  • Wikis, von denen erwartet wird, dass sie über 200 Seiten hinauswachsen
  • Bereiche mit vielfältigen Abfragetypen
  • Situationen, die hohen Recall erfordern
  • Multi-modale Wissensdatenbanken

Wann NICHT zu verwenden

  • Kleine Wikis (<100 Seiten) - index.md ist ausreichend
  • Einfache, gleichmäßige Inhalte
  • Situationen, in denen die Implementierungskomplexität nicht gerechtfertigt ist

Verwandte Concepts

Siehe auch

Beziehungen