Quickstart
This guide walks through a complete end-to-end example: loading an LLM, wiring up a retriever, running the pipeline, and inspecting the resulting schema.
Prerequisites
Set your LLM API key:
See Installation → Environment variables for the full list of variables (e.g. for the Weaviate retriever) and how to load them from a .env file.
Step 1 — Set up the LLM
SchemaGenerator accepts any LangChain BaseChatModel. Here we use OpenAI:
You are not tied to OpenAI: any OpenAI-compatible endpoint works, including a local Ollama/vLLM server or a LiteLLM proxy in front of another provider. See Configuration for the setup.
Step 2 — Set up a retriever
The data-grounded refinement stage needs a document retriever. Use the built-in HuggingFace adapter or implement your own.
from schematize.retrieval.huggingface import HuggingFaceRetriever
retriever = HuggingFaceRetriever(
dataset_name="JuDDGES/pl-court-raw",
text_column="text",
max_documents=10_000,
index_path=".cache/court-index", # cached after first build
)
The index is built lazily on first call and saved to index_path. Subsequent runs load from disk.
Step 3 — Load prompts
Bundled prompts are available for English and Polish, for tax and law domains:
from schematize import load_prompts
prompts = load_prompts(language="en", system_type="law")
# language: "en" | "pl"
# system_type: "law" | "tax"
load_prompts returns a flat dict of named prompt strings. Wrap it in SchemaGeneratorPrompts before passing to SchemaGenerator.
Step 4 — Build the generator
from schematize import SchemaGenerator, SchemaGeneratorPrompts
generator = SchemaGenerator(
llm=llm,
retriever=retriever,
prompts=SchemaGeneratorPrompts(**prompts),
max_refinement_rounds=3,
min_refinement_rounds=1,
max_data_refinement_rounds=2,
min_data_refinement_rounds=1,
data_assessment_top_k=20,
data_assessment_num_examples=3,
)
Step 5 — Run the pipeline
state = generator.stream_graph_updates(
"Study personal-rights violations in civil cases and assess their severity."
)
The pipeline will:
- Ask clarifying questions and wait for your answers (terminal input)
- Generate a search query
- Draft and refine the schema against quality criteria
- Retrieve documents and refine against real data
- Summarize the process and open a chat for final edits
Step 6 — Inspect the result
schema = state["current_schema"]
print(schema)
# {'fields': [{'name': 'violation_type', 'type_': 'enum', ...}, ...]}
Convert the schema to a Pydantic model for use in extraction:
from schematize import SchemaFields, DynamicModelFactory
spec = SchemaFields(**schema)
model_cls = DynamicModelFactory()(spec)
print(model_cls.model_fields)
Skipping pipeline stages
For faster iteration you can skip stages that don't apply to your use case:
generator = SchemaGenerator(
llm=llm,
retriever=retriever,
prompts=SchemaGeneratorPrompts(**prompts),
skip_problem_definition=True, # no clarification dialogue
skip_refinement=False,
skip_data_grounded=True, # no document retrieval
)
When skip_problem_definition=True, the user input is used directly as the problem definition.
Running from the command line
The steps above can be replaced with a single command if you install the [scripts] extra:
You'll be prompted to describe the schema to generate, then the same interactive pipeline from
Step 5 runs in your terminal. The resulting state is written to
state.json.
See the CLI reference for the full list of options (retriever selection, refinement rounds, output path, etc.).