Normal Behaviour Worker
Standalone normal behaviour worker: builds a description of what a typical session in a semantic
group looks like — an envelope for each metric, a centroid and representative example for the text
output, and a consensus or medoid execution graph — and writes a NormalBehaviourReport per layer.
Source: workers/normal-behaviour-worker
!!! info "Relationship to the analysis worker"
The default docker-compose deployment runs the combined
analysis worker, which embeds this same logic alongside anomaly detection
and consistency. This standalone worker exists for deployments that want to scale or schedule
normal-behaviour analysis separately (for example as an Argo Workflows step).
Where anomaly detection answers which sessions are unusual, this worker answers what does usual look like — a profile you can show a user, or compare a future session against.
Input
NormalBehaviourInputMessage:
{
"session_id": "session-abc",
"group_id": "group-xyz",
"group_hash": "abc123",
"sessions": [
{
"session_id": "session-abc",
"metrics": {"Cost": 1.0},
"output_content": "...",
"output_embedding": [0.1, 0.2],
"execution_graph": {"nodes": [], "edges": []}
}
]
}
| Field | Required | Notes |
|---|---|---|
session_id |
yes | Used for the idempotency pre-check and the group-membership guard |
group_id |
yes | The SemanticGroup to analyse |
group_hash |
no | Defaults to "", which disables the hash guard |
sessions |
no | Defaults to [] |
sessions selects the data path:
- Non-empty — analysed as-is; Neo4j is never read.
- Empty — needs a live DB connection. The worker runs
analysis_pre_check(group_id, "NormalBehaviourReport", group_hash, session_id), loads the group withget_analysis_data_for_semantic_group, and drops the message ifsession_idis not actually a member of the group.
Warning: Two behavioural differences from the other two workers -
--embedding-modelis checked first, beforesessions. Anomaly detection and consistency only require it on the database path; here a message is dropped without an embedding model even when it carries inline sessions and never touches Neo4j. - No hierarchical-grouping lock. This worker does not callwait_for_hierarchical_grouping_unlock(), so it can analyse a group while the hierarchical grouping worker is rewriting group membership. Thegroup_hashcheck insideanalysis_pre_checkis the only protection against a stale group.
Output
Queue output: none. The CLI hardcodes output_queue=[], so this worker is terminal.
Graph output: one NormalBehaviourReport per layer, linked from the SemanticGroup by
hasNormalBehaviourReport:
| Property | Value |
|---|---|
id |
sha256(group_id + layer + metric_name) |
dataType |
text, graph, or metric |
rawResult |
JSON dump of the full result object for that layer |
centroid |
JSON: the mean vector, median value, or consensus graph |
representativeSample |
JSON: the real data point closest to the centroid, or "" |
representativeProcessedSample |
For text, the actual output string of that session |
aboutMetric |
Set on metric reports only, linking to the Metric node |
The distinction between centroid and representativeSample matters: the centroid is usually
synthetic (an average embedding corresponds to no real session), while the representative sample
is an actual session you can show to a user.
File output (run-once): NormalBehaviourOutputMessage, echoing the session data it analysed
plus a normal_behaviour array:
{
"session_id": "session-abc",
"group_id": "group-xyz",
"sessions": [{"session_id": "session-abc", "...": "..."}],
"normal_behaviour": [
{
"normal_behaviour": {
"centroid": 1.2,
"std": 0.3,
"lower": 0.6,
"upper": 1.8,
"representative_sample": 1.2,
"statistic": "gaussian"
},
"session_ids": ["session-abc", "session-def"],
"layer": "metric",
"metadata": {"statistic": "gaussian", "metric": "Cost"}
}
]
}
Position in the pipeline
grouping-worker / hierarchical-grouping-worker
│
└──► new_session_to_normal_behaviour ──► normal-behaviour-worker ──► Neo4j
(NormalBehaviourReport)
| Input queue | new_session_to_normal_behaviour (NB_INPUT_QUEUE) |
| Output queue | none |
| CLI | normal-behaviour-worker-cli |
Warning: No producer publishes to this queue in the default deployment In
docker-compose.ymlthe grouping workers publish tonew_session_to_analysis, which the combined analysis worker consumes. Running this worker means either repointing a producer atnew_session_to_normal_behaviouror driving it in run-once mode.
Quick start
# Liveness check
uv run --directory workers/normal-behaviour-worker normal-behaviour-worker-cli --test
Queue mode:
normal-behaviour-worker-cli \
--rabbitmq-url amqp://<user>:<password>@rabbitmq:5672/ \
--input-queue new_session_to_normal_behaviour \
--embedding-model azure/text-embedding-3-small
Run-once against a file — note that --embedding-model is required even with inline sessions:
normal-behaviour-worker-cli \
--run-once \
--embedding-model azure/text-embedding-3-small \
--input /tmp/group.json \
--output /tmp/normal-behaviour.json
--run-once raises a usage error unless both --input and --output are given.
Build a wheel:
uv build workers/normal-behaviour-worker
Configuration
Copy .env.template to .env, or pass flags.
| Environment variable | CLI flag | Description |
|---|---|---|
RABBITMQ_URL |
--rabbitmq-url |
Full connection URL |
NB_INPUT_QUEUE |
--input-queue |
Input queue name |
NB_FEEDBACK_QUEUE |
--feedback-queue |
Optional feedback queue |
EMBEDDING_MODEL |
--embedding-model |
Always required |
NB_CONFIG_FILE |
--config-file |
YAML layer config |
| — | --max-sessions |
Stop after N messages (-1 = unlimited) |
NEO4J_URI, NEO4J_USERNAME, NEO4J_PASSWORD, NEO4J_DB |
(API-managed) | Read by the API DAL, not by CLI flags |
The YAML config selects which layers run and which statistic each uses. Omitting it is equivalent to:
text:
statistic: gaussian
graph:
statistic: consensus
metric:
statistic: gaussian
The graph layer accepts two extra keys:
graph:
statistic: consensus
majority_threshold: 0.5 # fraction of graphs an edge must appear in
get_closest_sample: false # skip the expensive nearest-real-graph search
An empty layer entry falls back to that layer's default; a layer left out of the file entirely is skipped.
Methodology
Source: dem/src/dem/normal_behaviour/.
Every layer produces the same conceptual triple:
sessions in the group
│
├─ centroid ──► the middle of the group
├─ envelope ──► the region counted as "normal"
└─ representative sample ──► the real session closest to the centroid
What differs per layer is the geometry: metrics live on a line, embeddings in a high-dimensional vector space, graphs in a space with no coordinates at all.
Metric layer
MetricNormalBehaviour (normal_metric.py) works on one metric at a time, over the group's
values for it. Three statistics are available.
gaussian (default)
Assumes the values are roughly normal and takes a k-sigma band, with k = 2.0:
centroid = mean(values)
std = standard_deviation(values)
lower = centroid - 2.0 * std
upper = centroid + 2.0 * std
For a truly normal distribution this covers about 95% of values. It is cheap and interpretable but it is dragged around by outliers, and metrics like latency or cost are usually right-skewed rather than normal — which is what the next option is for.
quantiles
Distribution-free: read the bounds straight off the empirical distribution.
lower = 5th percentile
upper = 95th percentile
centroid = median
The centroid is the median rather than the mean, so a single extreme session cannot move it. Prefer this for skewed or heavy-tailed metrics.
var_based
A narrower one-sigma band, with mean, median and variance all reported:
centroid = mean(values)
lower = centroid - std
upper = centroid + std
In all three cases representative_sample is set to the centroid. Because metrics are scalars the
centroid is itself a valid metric value, so no nearest-real-point search is needed. The statistic
field on the result is overwritten with the configured name (gaussian, quantiles,
var_based), so the internal detail of which k or which quantiles were used does not survive
into the report.
Text layer
TextNormalBehaviour (normal_textual.py) works on the matrix of output_embedding vectors, of
shape (n_sessions, embedding_dim). Six statistics are available.
gaussian, quantiles and var_based behave exactly as for metrics but are applied
independently per embedding dimension, so upper and lower are vectors describing an
axis-aligned box. This ignores correlation between dimensions, which for embeddings is
substantial — the box is a loose approximation of the true region.
centroid computes the mean vector only, with no envelope.
density
Fits a kernel density estimate over the embeddings and treats the low-density tail as abnormal:
KernelDensity(kernel="gaussian", bandwidth=1.0)
log_dens = kde.score_samples(embeddings)
threshold = quantile(log_dens, 1 - 0.95) # the 5th percentile of log-density
A session is normal when its log-density is above the threshold, so the "normal" region follows the
actual shape of the data instead of a box — it handles multi-modal groups, where sessions cluster
around two or three distinct kinds of output, which the Gaussian box cannot. The threshold is
stored under other_info["density_threshold"].
The fixed bandwidth=1.0 is the thing to watch: for high-dimensional normalised embeddings, where
typical distances are small, this heavily over-smooths and the estimate flattens out.
ellipse
The correlation-aware version of the Gaussian box. It estimates the full covariance matrix and uses squared Mahalanobis distance with a chi-squared cutoff:
d(x) = (x - mu)^T * Sigma^-1 * (x - mu)
threshold = chi2.ppf(0.95, df = n_dimensions)
normal <=> d(x) <= threshold
Under a multivariate normal assumption, d(x) follows a chi-squared distribution with n_dim
degrees of freedom, so the cutoff is a genuine 95% confidence ellipsoid rather than a per-axis
approximation.
Note: The ellipse parameters are not persisted
_calculate_confidence_ellipsepassescovarianceandmahalanobis_thresholdintoNormalBehaviourResultText, but neither is a declared field on that model, so Pydantic drops them. The report keeps the centroid and the representative sample; the ellipsoid itself cannot be reconstructed from what is stored. Usedensity, which stores its threshold in the declaredother_infofield, if you need the cutoff persisted.
Representative sample
Whatever the statistic, the final step is _find_closest, a plain nearest-neighbour search in
Euclidean distance:
index = argmin over i of || embeddings[i] - centroid ||
representative_sample becomes that session's embedding and representative_processed_sample its
actual output_content — the real text that best stands in for the group. As on the metric layer,
the statistic field is then overwritten with the configured name, so internal labels like
density_kde_0.95 or confidence_ellipsoid_0.95 never reach the report.
Graph layer
GraphNormalBehaviour (normal_graph.py). Graphs have no coordinates, so neither "mean" nor
"distance to the mean" is available directly. Two strategies are offered.
consensus (default)
Builds a synthetic graph out of the edges that most sessions agree on:
for each edge e:
count(e) = number of graphs containing e
keep e <=> count(e) > n_graphs * majority_threshold # default 0.5
nodes = endpoints of the kept edges
The result is the "majority route" through the system: the steps that happen on most runs, with one-off detours stripped out. Nodes are derived from the surviving edges rather than voted on separately, so a node that appears in every graph but never on a majority edge does not make it into the consensus.
Because that consensus graph may not match any real session, get_closest_sample (default True)
then searches for the actual graph nearest to it by networkx.graph_edit_distance and stores it as
the representative sample.
medoid
Picks the real graph that is most central — the one with the smallest total distance to all others. It builds the full pairwise edit-distance matrix and takes the row with the minimum sum:
D[i][j] = graph_edit_distance(G_i, G_j)
medoid = argmin over i of sum over j of D[i][j]
Graph edit distance is the minimum number of node and edge insertions, deletions and substitutions
needed to turn one graph into the other. Unlike consensus, the answer is always a real execution
graph, so centroid and representative_sample are the same graph; other_info additionally
carries the medoid's index and its distance to every other graph.
Warning: Graph edit distance is expensive Computing it is NP-hard;
networkxsearches for an optimal edit path and has no built-in timeout.medoidneedsn * (n-1) / 2of these, so it becomes impractical well before groups get large, andconsensuswith the defaultget_closest_sample: truestill needsnof them. For large groups setget_closest_sample: false— you keep the consensus graph and lose only the representative sample.
Caveats
- At least two sessions are required.
process_groupreturns an empty list for groups of size 0 or 1, and the worker still returnsTrue, so the message is acked with no reports written. - The idempotency pre-check only runs on the database path. Passing inline
sessionsbypassesanalysis_pre_check, so reports are recomputed and re-MERGEd on every call. - Zero reports still counts as success. When no layer produces a report the worker returns
Truewithout replacingoutput_messages, so run-once mode writes the input message back out unchanged rather than aNormalBehaviourOutputMessage. - A layer is skipped silently when its data is missing. Sessions with no
output_embedding, noexecution_graphor nometricsare excluded from the corresponding layer, and if that leaves nothing the layer produces no report at all.