Observability with MLflow
Synalinks provides built-in observability through MLflow, enabling you to trace and monitor your LM programs in production.
Overview
The observability system automatically creates spans for each module call, capturing:
- Inputs and outputs of each module; language model and decision model spans carry the chat messages and tools in OpenAI format, so the MLflow UI renders the conversation
- Token usage and cost on language model and decision model spans, using
MLflow's standard
mlflow.chat.tokenUsage,mlflow.llm.cost,mlflow.llm.modelandmlflow.llm.providerattributes, so the trace-level token and cost roll-ups work - Duration and success/failure status of each call, including the model
calls that failed every retry (they return
Noneinstead of raising) - Parent-child relationships between nested module calls
- Phase tags: every trace is tagged
synalinks.programandsynalinks.phase(inference,rewardoroptimizer), so the calls made by rewards and optimizers during training can be filtered out of production views - Linked prompts: calls of a module whose prompt was registered in the Prompt Registry link the trace to that prompt version
Quick Start
Enable Observability
Important: You must call
enable_observability()BEFORE creating any modules. Hooks are registered when modules are instantiated, so enabling observability after module creation will not trace those modules.
import synalinks
# Enable FIRST, before creating any modules
synalinks.enable_observability(
tracking_uri="http://localhost:5000", experiment_name="my_experiment"
)
# Now create your modules - they will be automatically traced
inputs = synalinks.Input(data_model=Question)
outputs = await synalinks.Generator(...)(inputs)
Once enabled, all module calls in your program will be automatically traced.
Example Usage
import asyncio
import synalinks
# Enable observability before creating your program
synalinks.enable_observability(
tracking_uri="http://localhost:5000", experiment_name="question_answering"
)
class Question(synalinks.DataModel):
question: str = synalinks.Field(description="The user's question")
class Answer(synalinks.DataModel):
answer: str = synalinks.Field(description="The answer to the question")
async def main():
language_model = synalinks.LanguageModel(model="gemini/gemini-3.1-flash-lite-preview")
# Create a simple question-answering program
inputs = synalinks.Input(data_model=Question)
outputs = await synalinks.Generator(
data_model=Answer,
language_model=language_model,
)(inputs)
program = synalinks.Program(
inputs=inputs,
outputs=outputs,
name="qa_program",
description="A simple QA program",
)
# Run the program - traces will be sent to MLflow
result = await program(Question(question="What is the capital of France?"))
if result:
print(result.prettify_json())
if __name__ == "__main__":
asyncio.run(main())
Running MLflow with Docker
Using Docker
Run MLflow tracking server locally with artifact proxying enabled:
docker run -d \
--name mlflow \
-p 5000:5000 \
-v mlflow-data:/mlflow \
ghcr.io/mlflow/mlflow:latest \
mlflow server \
--host 0.0.0.0 \
--port 5000 \
--backend-store-uri sqlite:///mlflow/mlflow.db \
--default-artifact-root mlflow-artifacts:/ \
--serve-artifacts \
--artifacts-destination /mlflow/artifacts
Important flags:
--serve-artifacts: Enables the MLflow server to proxy artifact uploads from clients--default-artifact-root mlflow-artifacts:/: Tells clients to use the server as an artifact proxy--artifacts-destination /mlflow/artifacts: Where the server stores artifacts on disk
Then configure Synalinks to use it:
import synalinks
synalinks.enable_observability(
tracking_uri="http://localhost:5000", experiment_name="synalinks_traces"
)
Using Docker Compose
For a more complete setup with persistent storage, create a docker-compose.yml:
services:
mlflow:
image: ghcr.io/mlflow/mlflow:latest
container_name: mlflow
ports:
- "5000:5000"
volumes:
- ./mlflow-data:/mlflow
command: >
mlflow server
--host 0.0.0.0
--port 5000
--backend-store-uri sqlite:///mlflow/mlflow.db
--default-artifact-root mlflow-artifacts:/
--serve-artifacts
--artifacts-destination /mlflow/artifacts
restart: unless-stopped
Start the services:
Access the MLflow UI at http://localhost:5000.
Production Setup with PostgreSQL
For production deployments, use PostgreSQL as the backend store:
services:
postgres:
image: postgres:16-alpine
container_name: mlflow-postgres
environment:
POSTGRES_USER: mlflow
POSTGRES_PASSWORD: mlflow
POSTGRES_DB: mlflow
volumes:
- postgres-data:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U mlflow"]
interval: 5s
timeout: 5s
retries: 5
restart: unless-stopped
mlflow:
image: ghcr.io/mlflow/mlflow:latest
container_name: mlflow
depends_on:
postgres:
condition: service_healthy
ports:
- "5000:5000"
volumes:
- mlflow-artifacts:/mlflow/artifacts
environment:
MLFLOW_BACKEND_STORE_URI: postgresql://mlflow:mlflow@postgres:5432/mlflow
command: >
mlflow server
--host 0.0.0.0
--port 5000
--backend-store-uri postgresql://mlflow:mlflow@postgres:5432/mlflow
--default-artifact-root mlflow-artifacts:/
--serve-artifacts
--artifacts-destination /mlflow/artifacts
restart: unless-stopped
volumes:
postgres-data:
mlflow-artifacts:
Understanding Traces
When you run a Synalinks program with observability enabled, MLflow captures detailed traces.
Span Types
Synalinks automatically categorizes spans based on module type for better visualization in MLflow:
| Module | Span Type |
|---|---|
LanguageModel, DecisionModel |
CHAT_MODEL |
EmbeddingModel |
EMBEDDING |
FunctionCallingAgent |
AGENT |
EmbedKnowledge, RetrieveKnowledge, UpdateKnowledge |
RETRIEVER |
Tool |
TOOL |
Other modules (Generator, ChainOfThought, SelfCritique...) |
CHAIN |
A Generator is a CHAIN span wrapping the CHAT_MODEL span of the model it
calls. A DecisionModel is called like a LanguageModel, so its calls
are traced the same way.
Span Attributes
Each span includes these attributes:
| Attribute | Description |
|---|---|
synalinks.call_id |
Unique identifier for this call |
synalinks.parent_call_id |
ID of the parent call (for nested modules) |
synalinks.module |
Module class name (e.g., Generator) |
synalinks.module_name |
Custom name given to the module |
synalinks.module_description |
Module description |
synalinks.is_symbolic |
Whether the call was symbolic (graph building) |
synalinks.duration |
Call duration in seconds |
synalinks.success |
Whether the call succeeded |
Model spans (CHAT_MODEL) also carry:
| Attribute | Description |
|---|---|
mlflow.llm.model / mlflow.llm.provider |
The model requested (e.g. openai/gpt-4o-mini) and its provider |
mlflow.chat.tokenUsage |
Input, output and total tokens of the call |
mlflow.llm.cost |
Cost of the call in USD (total_cost), when known |
synalinks.cache_hit |
The response came from the model's on-disk cache |
synalinks.fallback |
Every attempt failed and the fallback model answered |
Decision model spans add the model that answered and its raw answers:
| Attribute / Output | Description |
|---|---|
synalinks.response_model |
The versioned model that answered (e.g. typesafe/jev-1.13.0 for jev-latest), which the cost is based on |
answers (span outputs) |
The raw answers, with the probabilities and confidence the output values leave out |
Exception Events
When a module call fails, the span automatically records an exception event with:
- exception.type: The exception class name
- exception.message: The exception message
A language or decision model call that fails every retry returns None rather
than raising, so the program can carry on. Its span is still marked as failed
(ERROR status, synalinks.success false, exception event with the last error).
When a fallback model answers instead, the failed call's span only records the
failure and synalinks.fallback, and the fallback call gets its own child span
with its tokens and cost, so they are counted once.
Users and Sessions
MLflow can associate traces with a user and a chat session, which is useful for
analyzing multi-turn conversations: you can then filter the traces of one
session in the UI and inspect what happened at each turn. MLflow reads two
reserved metadata keys for that, mlflow.trace.user and mlflow.trace.session.
Synalinks creates its spans outside of MLflow's fluent context (nested module
calls are linked explicitly), so mlflow.update_current_trace() does not see
them. Use synalinks.trace_context() instead: every trace started inside the
block carries the given user and session.
import synalinks
synalinks.enable_observability()
# ... build your program ...
with synalinks.trace_context(user_id="user-123", session_id="session-123"):
result = await program(inputs)
The context is stored in a contextvars.ContextVar, so it is safe to set per
request in an async server: concurrent requests each see their own value, and
the tasks spawned inside the block (parallel branches, agent tool calls)
inherit it. Nested blocks merge with the enclosing one, the innermost value
winning.
@app.post("/chat")
async def chat(request: ChatRequest):
with synalinks.trace_context(user_id=request.user_id, session_id=request.session_id):
return await program(request.messages)
You can also attach free-form trace metadata (immutable once the trace is
logged) and tags (editable afterwards in the MLflow UI):
with synalinks.trace_context(
user_id="user-123",
session_id="session-123",
metadata={"turn": "3"},
tags={"env": "production"},
):
result = await program(inputs)
Viewing Traces
- Open MLflow UI at
http://localhost:5000 - Navigate to your experiment
- Click on a run to see detailed traces
- Use the trace view to explore the call hierarchy
Configuration Options
Environment Variables
You can also configure MLflow using environment variables:
Then in your code:
import synalinks
# Will use MLFLOW_TRACKING_URI from environment
synalinks.enable_observability(experiment_name="my_experiment")
Direct Monitor Hook
For fine-grained control, you can create a Monitor hook directly:
import synalinks
monitor = synalinks.hooks.Monitor(
tracking_uri="http://localhost:5000", experiment_name="custom_experiment"
)
# Add to a specific module
generator = synalinks.Generator(data_model=Answer, hooks=[monitor])
Training Metrics and Artifacts
The Monitor callback logs each fit() run to MLflow. When enable_observability()
was called it is added to every fit() and evaluate() automatically; pass one in
callbacks=[...] to customize it. A training run holds:
- Metrics, one point per epoch: the training metrics, the
val_*validation metrics, the spend of the epoch (epoch_cost,epoch_tokens,epoch_calls, andepoch_cost_inference/epoch_cost_reward/epoch_cost_optimizerwhen non-zero), the running spend of the fit (fit_cost,fit_tokens), the all-time in-process spend (total_cost,total_tokens),epoch_duration_s, andfit_duration_sat the end. Withlog_batch_metrics=True, batch metrics are logged astrain_batch_*andval_batch_*with their own step counters. - Params: the fit arguments (epochs, batch and minibatch size, validation split and
frequency, dataset sizes), the compile configuration (reward, metrics, optimizer and
its config), the program (name, class, number of modules and trainable variables)
and every language or embedding model reachable from the program
(
lm.<name>.model, api base, sampling settings). - Datasets: the train and validation sets as MLflow dataset inputs of the run, one
row per sample with JSON
inputsandexpectationscolumns. - The program plot, the program model and its prompts, described below.
Basic Usage
import synalinks
# Create the monitor callback
monitor = synalinks.callbacks.Monitor(
tracking_uri="http://localhost:5000",
experiment_name="training_experiment",
run_name="my_training_run",
log_program_plot=True, # Save program visualization as artifact
)
# Use during training
program.fit(x=train_inputs, y=train_labels, epochs=10, callbacks=[monitor])
Program Plot Artifact
When log_program_plot=True (the default), the Monitor callback automatically saves
a visualization of your program architecture as an MLflow artifact at the start of training.
The plot is saved under program_plots/ in the artifacts folder and includes:
- Module names and types
- Input/output schemas
- Trainable status of each module
You can view the program plot in the MLflow UI under the "Artifacts" tab of your run.
Program Model
When log_program_model=True (the default), the Monitor callback logs the whole
program (architecture and trained variables, the JSON written by Program.save()) as
an MLflow pyfunc model wrapped in synalinks.callbacks.SynalinksProgramModel.
A new model is logged every time val_reward improves (reward without validation
data, the end of training when neither is available). Each is an MLflow 3
LoggedModel version of type synalinks_program, linked to the training run and
carrying the epoch metrics logged at that step, so the Versions view of the experiment
compares them side by side; the best one is tagged synalinks.best. Every logged
model is also registered in the Model Registry under the program's name, so
models:/<program name>/<version> and registry aliases work out of the box.
- The model can be loaded and served anywhere
synalinksis installed:
import mlflow
model = mlflow.pyfunc.load_model(f"models:/{model_id}")
model.predict({"question": "What is the capital of France?"})
# [{"answer": "Paris"}]
or with mlflow models serve -m models:/<model_id>. The model's signature is derived
from the program's input and output JSON schemas (inferred from the first training
sample when the schema cannot be expressed). Custom DataModel, Module or Program
subclasses are resolved by Program.load() from their import path, so the code
declaring them only needs to be importable where the model is loaded.
The model id is written into the saved program (program._mlflow_model_id, persisted
by Program.save() under the mlflow key). A later evaluate() of that program, even
in another process after Program.load(), links its metrics and traces to that model
version and nests its run under the training run.
Prompt Registry
The prompts of the program's trainable
modules (Generator, ChainOfThought, ... every module holding an Instructions
variable) are registered in the MLflow Prompt Registry under the name
<program>.<module>. Each registered version is a chat prompt: the system turn is the
module's prompt rendered from its optimized instructions and few-shot examples, the
user turn holds the {{ inputs }} template variable, response_format is the module's
output JSON schema and model_config its language model settings, so the version can
be replayed in the MLflow Playground and evaluated with mlflow.genai.evaluate().
A new version is created at every epoch where the rendered prompt changed. Its commit
message and val_reward tag record the validation reward of that epoch (the state the
validation scored is the one registered), along with the epoch, run id, program, module
and optimizer. At the end of training the best-scoring version receives the best
alias:
Versions are linked to the training run, to the logged model (prompts= of the model,
shown on the model page) and to every later trace of the module (mlflow.linkedPrompts).
Callback Parameters
| Parameter | Default | Description |
|---|---|---|
experiment_name |
Program name | MLflow experiment name |
run_name |
Program name | MLflow run name, suffixed _train / _test |
tracking_uri |
Local ./mlruns |
MLflow tracking server URI |
log_batch_metrics |
False |
Log metrics at batch level |
log_epoch_metrics |
True |
Log metrics at epoch level |
log_program_plot |
True |
Save program visualization as artifact |
log_program_model |
True |
Log the program as a registered MLflow model, a new version per reward improvement |
tags |
{} |
Additional tags for the run |
run_id |
None |
Existing run to resume instead of starting a new one; the step counter continues after the last logged step |
resume |
False |
Look up a run named run_name in the experiment and resume it, creating it the first time |
log_assessments |
True |
Log each evaluated sample's reward as a feedback assessment on its trace |
Example with Full Configuration
import synalinks
monitor = synalinks.callbacks.Monitor(
tracking_uri="http://localhost:5000",
experiment_name="gsm8k_optimization",
run_name="chain_of_thought_v1",
log_batch_metrics=True,
log_epoch_metrics=True,
log_program_plot=True,
tags={"model": "gpt-4o-mini", "optimizer": "RandomFewShot", "dataset": "gsm8k"},
)
program.fit(x=train_questions, y=train_answers, epochs=5, callbacks=[monitor])
Evaluation Results Over Time
program.evaluate() with the Monitor callback logs the reward and metrics of that
evaluation to an MLflow run. Runs without a logged model are listed in the experiment's
Evaluation runs tab (/#/experiments/<id>/evaluation-runs), where the charts view
compares the metrics of every evaluation. The run is tagged mlflow.runType =
genai_evaluate, like the runs created by mlflow.genai.evaluate(), and the traces of
the evaluated calls are linked to it. Each evaluated sample's reward is also logged as
a reward feedback assessment on the sample's trace, with the ground truth as an
expected_output expectation, so the run's traces can be sorted and filtered by score
(disable with log_assessments=False). The run also logs eval_cost, eval_tokens,
eval_duration_s and the evaluated dataset.
When the evaluated program was trained with a Monitor (or loaded from a program saved
after such a training), the evaluation run is nested under the training run
(mlflow.parentRunId) and its metrics and traces are linked to the logged model
version, so the Versions view shows evaluation results per model version.
When synalinks.enable_observability() was called, a Monitor callback created without
experiment_name uses that same experiment, so the evaluation runs appear next to the
traces:
import synalinks
synalinks.enable_observability(
tracking_uri="http://localhost:5000", experiment_name="my_app"
)
# ... build and compile your program ...
monitor = synalinks.callbacks.Monitor(run_name="eval")
await program.evaluate(x=test_x, y=test_y, callbacks=[monitor])
Each evaluate() call creates one run, so the Evaluation runs charts show one point per
evaluation. To draw the metrics as a curve inside a single run instead, like the epoch
curves of fit(), resume the same run on every evaluation with resume=True: the
callback looks up the run named run_name in the experiment (creating it the first
time), and logs each evaluation at the next step. This works across processes, so a
scheduled evaluation appends one point per execution:
monitor = synalinks.callbacks.Monitor(run_name="nightly_eval", resume=True)
await program.evaluate(x=test_x, y=test_y, callbacks=[monitor])
In the MLflow UI, open the run and plot the metric against its step, or against wall
time to see the results over time. Pass run_id instead of resume to resume a
specific run; monitor.run_id holds the id of the run in use.
Combining Tracing with Training
When using both enable_observability() and the Monitor callback for training:
-
During program building (symbolic calls): Traces go to the experiment specified in
enable_observability() -
During training (
fit()): Traces are associated with the training run and go to the experiment of theMonitorcallback. When the callback has noexperiment_name, that is the experiment ofenable_observability(), so traces and runs stay together; give the callback its ownexperiment_nameto keep them apart.
Full Example
import synalinks
# Enable tracing for all module calls
synalinks.enable_observability(
tracking_uri="http://localhost:5000",
experiment_name="synalinks_traces", # Traces during setup go here
)
# Create your program (symbolic traces created here)
inputs = synalinks.Input(data_model=Question)
outputs = await synalinks.Generator(
data_model=Answer,
language_model=language_model,
)(inputs)
program = synalinks.Program(inputs=inputs, outputs=outputs, name="my_program")
# Create Monitor callback for training
monitor = synalinks.callbacks.Monitor(
tracking_uri="http://localhost:5000",
experiment_name="training_runs", # Training metrics + traces go here
run_name="experiment_v1",
)
# Train - traces during fit() are associated with the training run
program.compile(reward=reward, optimizer=optimizer)
await program.fit(x=train_x, y=train_y, epochs=5, callbacks=[monitor])
After training, you'll have: - synalinks_traces experiment: Setup traces (symbolic module calls) - training_runs experiment: Training run with metrics, artifacts, and execution traces
Best Practices
- Enable observability early in your script, before creating any modules
- Use meaningful experiment names to organize your traces by project or feature
- Use persistent storage (PostgreSQL) for production deployments
- Set up retention policies to manage storage for long-running applications
Troubleshooting
No traces being created
If you don't see any traces in MLflow:
-
Check call order: Ensure
enable_observability()is called before creating any modules -
Verify observability is enabled: Check with
synalinks.is_observability_enabled() -
Check the correct experiment: During training, traces go to the training experiment, not the observability experiment
MLflow not receiving traces
- Verify the MLflow server is running:
curl http://localhost:5000/health - Check the tracking URI is correct
- Ensure
mlflowpackage is installed:pip install mlflow
Artifacts not showing in MLflow UI
If artifacts are uploaded but don't appear in the MLflow UI:
- Check server configuration: Ensure the MLflow server is started with
--serve-artifactsflag - Verify artifact root: The server must use
--default-artifact-root mlflow-artifacts:/for remote clients - Check permissions: The server needs write access to
--artifacts-destinationpath
Correct server configuration:
mlflow server \
--serve-artifacts \
--default-artifact-root mlflow-artifacts:/ \
--artifacts-destination /mlflow/artifacts
Common mistake - Missing --serve-artifacts causes clients to try writing directly to the server's local filesystem, resulting in permission errors like:
Missing cost information
Cost tracking requires the language model to return usage information. Ensure your LLM provider supports this feature.
Decision models are billed on input tokens only, at a price known per versioned
model. For a model missing from the built-in price table, pass cost_per_token
to the DecisionModel.