Skip to main content
POST

Overview

Run retrieval-augmented generation (RAG) on the local store:
  1. Search using the provided query_vector (same dimension as the store).
  2. Build a prompt from the top matching passages.
  3. Call Ollama on the host with fixed model qwen2.5:0.5b-instruct.
If no passages pass search (empty store or nothing above threshold in kiosk mode), the server does not call the LLM and returns: I don't have enough information to answer that question.
Requires Ollama running on the host. Start the stack with moorcheh-edge up --with-llm (use --skip-ollama for search-only). The CLI and SDK embed query locally before calling this endpoint.

Request body

query
string
required
Original question text (included in the LLM prompt and echoed in the response).
query_vector
array
required
JSON array of floats used for similarity search. Length must match the store dimension (768 for text stores).
top_k
number
default:"5"
Number of passages to retrieve for context. Capped at 100.
threshold
number
default:"0"
Minimum search score when kiosk_mode is true.
kiosk_mode
boolean
default:"false"
When true, filters retrieved passages below threshold.
header_prompt
string
Optional system instruction (replaces the default RAG system prompt).
Optional instruction appended before the user question in the final user message.
chat_history
array
Prior turns: [{"role": "user"|"assistant", "content": "..."}].
temperature
number
default:"0.2"
LLM sampling temperature (0.0–2.0).

Response fields

Errors