16 KiB
KnowledgeFS Operator Manual
This manual is for people running KnowledgeFS in development, staging, or production. It complements the API reference and deployment guide with daily operating procedures, quality gates, incident response, and performance guardrails.
Operating Model
KnowledgeFS is split into independently observable services:
| Service | Responsibility |
|---|---|
| Admin Console | Human workflows, upload/evaluation dashboards, Retrieval Studio, trace diagnostics. |
| Hono API | Auth, ingestion, retrieval, KnowledgeFS, queries, evaluation routes, traces, MCP tools. |
| Database | Tenant-scoped metadata, generated artifacts, nodes, projections, traces, evaluation data. |
| Object storage | Raw uploaded document bytes. |
| Parser service | Unstructured-compatible parsing for complex document formats. |
| Queue runtime | Async document compilation, bulk jobs, cleanup, and research work when configured. |
| TypeScript compute | Pure bounded compute: chunking, token counting, RRF, packing, diff. |
The Admin Console must not bypass the Hono API for business data. The API is the security and tenant boundary.
Daily Health Checks
Run these at the start of each operating day and after each deployment:
curl -fsS "$API/health"
curl -fsS "$API/openapi.json" >/dev/null
pnpm eval:regression
For Standalone environments:
docker compose --env-file infra/local/.env -f infra/local/compose.yaml --profile apps ps
docker compose --env-file infra/local/.env.example -f infra/local/compose.yaml --profile apps config >/dev/null
Expected health:
- API returns healthy platform adapter status.
- Parser, embedding, LLM, reranker, object storage, database, cache, and job components are either healthy or explicitly marked unavailable for the environment.
- Retrieval regression gate passes recall, citation-hit, no-answer, citation accuracy, and faithfulness thresholds.
Release Checklist
Before promoting a release candidate:
pnpm install --frozen-lockfile
pnpm check
pnpm build
pnpm lint
pnpm compose:config
docker compose --env-file infra/local/.env.example -f infra/local/compose.yaml --profile apps config
pnpm docker:api:build
pnpm docker:api:bundle-smoke
git diff --check
docker:api:bundle-smoke deliberately starts the built API bundle with NODE_ENV=test. It proves
that the container can boot, serve /health, and report components.compute === true; it does not
exercise production fail-closed startup, database repositories, durable compilation, object
storage, or providers. Production promotion still requires the deployed/Compose-backed health and
tenant-scoped upload/query checks below. The legacy docker:api:http-smoke command is only an alias
for this isolated check.
Confirm:
.harness/changescontains the change record for the release slice..harness/docs/TEMP-progress-document.mdrecords RED/GREEN verification and commit count.- TypeScript compute tests and coverage gates passed.
- Database migration drift check passed.
- The implementation commit count since the latest review checkpoint is below 10, or the mandatory health review has been completed.
Tenant And Auth Operations
Business routes require bearer auth. A valid subject includes subjectId, tenantId, and scopes.
Use scoped test tokens for smoke checks:
knowledge-spaces:readfor read-only checks.knowledge-spaces:writefor upload and mutation checks.knowledge-spaces:*only for trusted administrative smoke flows.
Operational rules:
- Never put
tenantIdin client requests expecting it to be trusted. - Treat cross-tenant 404s as expected behavior.
- Rotate
AUTH_JWT_SECRETor provider secrets through the environment secret manager, not through committed files. - Never log bearer tokens.
Ingestion Operations
Single-file ingestion:
- Create or choose a KnowledgeSpace.
- For a new space, select its Dify-managed
pluginId,provider, and embeddingmodelat creation time or withPUT /knowledge-spaces/{id}/embedding-profilebefore uploading data. KnowledgeFS sends that routing identity to Dify's inner model API; Dify resolves the workspace model instance and credentials. Do not configure or copy model credentials into KnowledgeFS. - Do not configure a vector dimension; it is observed from the selected model and persisted by the service. Select the profile before the first ingestion. Ingestion atomically freezes the profile, and any later change requires the reindex/publish workflow (even if that first upload subsequently fails).
- When rolling this admission-latch release into an existing cluster, drain older ingestion instances before enabling profile updates; older binaries do not stamp the latch.
- For a new space, select its Dify-managed
- Upload
multipart/form-datafieldfileto/knowledge-spaces/{id}/documents. - Check response:
201means synchronous MVP parsing completed.202means async compilation was queued.500withDocument parsing failedmeans the raw object and asset should remain for retry or diagnostics.
- Fetch
/knowledge-spaces/{id}/documents/{documentId}. - Fetch
/knowledge-spaces/{id}/documents/{documentId}/parse-artifacts/{version}when parsed.
Bulk ingestion:
- Keep file count and total byte size within configured limits.
- Use bulk upload when document compilation jobs are configured.
- Monitor
bulkJobIdand per-document status URLs. - If a bulk upload fails after object writes, cleanup is best-effort; inspect object storage for leftover keys under the tenant/space prefix.
Parser failure triage:
| Symptom | Likely Cause | Action |
|---|---|---|
400 upload error |
Missing multipart file or invalid body | Retry with file field. |
413 upload error |
File size or quota exceeded | Reduce file size or adjust quota after review. |
500 Document parsing failed |
Parser or artifact persistence failed | Check x-trace-id, parser component health, and artifact repository logs. |
Asset stuck pending |
Async compilation worker unavailable | Check queue runtime and job status. |
Asset failed |
Parser/job failure or status update after job start failure | Reindex after fixing dependency. |
Retrieval And Query Operations
Use /queries for user-facing retrieval plus generation. The endpoint streams SSE and records an answer trace.
The service has three retrieval pipelines and one optional public router:
- Fast runs ordinary dense + FTS hybrid recall, candidate fusion, and the configured final rerank.
- Research uses published Summary/Outline/PageIndex navigation. It does not run ordinary hybrid recall, Graph expansion, or the ordinary candidate reranker.
- Deep runs ordinary hybrid recall, adds permission-scoped Graph expansion, merges both candidate sets, and then runs one unified final rerank.
- An explicit
mode: "auto"asks the knowledge space's publishedreasoningModelthrough the Dify model runtime to choose one of those pipelines. Auto is not a fourth pipeline. OmittingmodeusesdefaultModedirectly, and explicit concrete modes bypass the router.
Auto routing is model-based; there is no CJK/language, query-length, word-count, or keyword
heuristic fallback. On timeout, provider failure, invalid structured output, or model-identity
mismatch, the request safely uses the published defaultMode. Treat repeated fallback decisions
as a reasoning-provider health signal, not as successful classifier behavior.
Operate with these checks:
- Record
x-trace-idfor HTTP/log/OTLP correlation,x-query-run-id(or SSEdata.traceId) for the durable AnswerTrace resource, andx-session-idfor session continuation. These IDs are not interchangeable. - Fetch
/queries/{traceId}withx-query-run-idor SSEdata.traceId, never the transportx-trace-id. - Use
/queries/{traceId}/evidence,/conflicts, and/missingfor bounded virtual evidence views. - Use the Admin trace comparison and failed query diagnostics panels to compare routing, recall candidates, filters, rerank changes, and evidence bundles.
- Inspect the persisted
query.routestep when diagnosing mode selection. It recordsrequestedMode, concreteresolvedMode,resolver(explicit,llm, orfallback), prompt version, bounded model/provider/usage provenance, duration, anddegradedplus a safe error class for fallback. It never contains the router prompt or raw model response and is not streamed as an SSE event. - For asynchronous Research jobs, an explicit Auto decision is made once against the frozen published profile during job creation. The concrete mode and bounded routing provenance are persisted; queue retries, lease recovery, and worker restarts must reuse that decision rather than invoke the classifier again.
- Before deploying this contract over a database that may contain unfinished legacy Research jobs
with
mode=auto, backfill each job to a reviewed concrete mode or cancel it. Workers fail closed on unresolved legacy Auto jobs because replaying the old heuristic would violate the frozen model/publication contract.
Do not place raw answer text, document chunks, prompts, JWTs, uploaded bytes, or AnswerTrace evidence text in operational logs/OTLP attributes unless an incident-specific data handling process authorizes it. AnswerTrace itself intentionally persists authorized evidence text inside its EvidenceBundle, so apply the same data-classification and access controls to trace storage.
Evaluation Operations
Evaluation quality is governed by:
- Golden question CRUD.
- Automatic question generation with human review.
- Human annotation workflow.
- Advanced metrics: context precision, relevance, faithfulness, citation accuracy.
- A/B retrieval strategy comparison.
- CI regression gate.
Routine flow:
- Capture production bad cases from failed traces.
- Review generated or captured questions before they enter the golden set.
- Add human annotations for answer correctness and evidence relevance.
- Run strategy comparisons against the same bounded golden set.
- Promote retrieval or prompt changes only when
pnpm eval:regressionpasses.
Regression gate failures:
| Failure | Meaning | Response |
|---|---|---|
totalQuestions below minQuestions |
Sample is too small to trust. | Restore or regenerate the evaluation report. |
recallAtK below minRecallAtK |
Retrieval missed expected evidence. | Inspect candidate ranking, filters, index freshness. |
citationHitRate below minCitationHitRate |
Citations do not cover expected evidence ids. | Inspect citation normalization and source locations. |
citationAccuracy below minCitationAccuracy |
Judge found unsupported or wrong citations. | Review answer/evidence alignment. |
faithfulnessScore below minFaithfulnessScore |
Judge found unsupported answer claims. | Review prompts, evidence packing, and generation model behavior. |
noAnswerRate exceeds maxNoAnswerRate |
System is abstaining too often. | Inspect retrieval thresholds and answerability classifier. |
KnowledgeFS Operations
KnowledgeFS routes provide bounded filesystem-like inspection:
ls,tree,findfor navigation.cat,stat,open_nodefor inspection.grepfor search.difffor version comparison.
Rules:
- Always supply explicit limits.
- Prefer
open_nodefor citation-ready node inspection. - Use
difffor troubleshooting stale or changed document versions. - Do not run ad hoc database scans to recreate KnowledgeFS views; use the bounded API or repository tools.
Storage And Retention
Object storage:
- Raw documents are stored under tenant/space/document prefixes.
- Object metadata includes asset id, KnowledgeSpace id, tenant id, hash, and uploader when available.
- Production should use S3-compatible object storage, MinIO, or R2. Bounded memory storage is for development only.
Retention:
- Use tenant-level and KnowledgeSpace-level retention policy routes to configure cleanup cutoffs. The policy is declarative: verify that retention workers are scheduled and monitor their job results, because PATCH does not synchronously delete retained data.
- Do not bulk-delete object prefixes manually unless the database cascade state has been reviewed.
- Use bulk delete APIs for bounded cascade tracking across assets, artifacts, nodes, projections, objects, and lifecycle records.
Performance Guardrails
Treat performance regressions as correctness failures:
- No unbounded list, dequeue, stream read, upload, provider response, or cache entry.
- Every database read path needs an explicit
maxRowsor route-level limit. - Avoid N+1 queries; prefer repository methods that join or batch required data.
- Cache keys must include tenant, subject or permission snapshot, strategy, model, and index versions where relevant.
- Queue, retention, in-memory fallback, and Admin diagnostic surfaces must keep explicit max sizes.
- Never add a hot path that fetches object storage bytes after upload when bytes are already in memory.
Incident Response
Use this order during production incidents:
- Identify blast radius: tenant, KnowledgeSpace, document ids, trace id, job id, or bulk job id.
- Check
/healthand component-level health. - Check recent deploy commit and
.harness/changesrecord. - Gather bounded evidence: trace, job status, document asset, parse artifact, KnowledgeFS
stat/open_node. - Stop the unsafe path:
- Disable traffic to Admin for UI-only bugs.
- Roll back API for ingestion, retrieval, auth, persistence, or queue bugs.
- Pause workers or queue consumers for runaway async work.
- Preserve data. Do not delete database rows or object prefixes without a recovery plan.
- Add a regression test or evaluation case before closing the incident.
Rollback Procedure
- Stop or shift traffic from the faulty service.
- Roll back API first for backend or data-path issues.
- Roll back Admin first only for UI-only issues.
- Keep database migrations in place unless a reviewed down-migration exists.
- Keep object storage data in place.
- Re-run smoke checks and
pnpm eval:regression. - Record the rollback in
.harness/changes.
Observability
Trace ids:
- Every response should include
x-trace-idfor transport correlation. Query streams additionally exposex-query-run-idas the durable Query/AnswerTrace identity. - Ingestion spans include bounded steps such as space lookup, upload read/hash, object put, asset create, parser parse, artifact create, status update, cleanup, and failure marking.
- Query traces record evidence, conflicts, missing evidence, and generation metadata.
- Query traces include
query.routeso operators can distinguish an explicit/default concrete selection, an LLM Auto selection, and a degraded Auto fallback without logging query content.
Safe attributes:
- Route, method, status, tenant id, subject id, low-cardinality error class, job id, trace id.
Forbidden attributes:
- JWTs.
- Raw file bytes.
- Full document text.
- Provider prompts or raw model responses.
- Secrets or object bodies.
Admin Console Workflows
Use the Admin Console for:
- Upload health and retrieval preview.
- Trace viewer and trace comparison.
- Evaluation dashboard.
- Retrieval Studio comparison.
- Golden question management and generated-question review.
- Human annotation workflow.
- Failed query diagnostics.
The Admin BFF is a thin proxy and must not become a second business API.
When To Escalate
Escalate before continuing feature work when:
- The 10-commit review cadence is due.
- Coverage drops below 90%.
pnpm check,pnpm build,pnpm lint, or Compose config fails.- A change needs production secrets, live database migration execution, or external provider account changes.
- A proposed fix requires deleting tenant data or object storage prefixes.