Integration

Observe, fix and optimise with Aspire

Send every service’s telemetry from your Aspire app to RepoQL. Your agent can then find what failed and what’s slow across the whole system, follow each one to the line of code, and prove a fix with a before-and-after query.

The route

  1. Connect the AppHost. Add one package and one line.
  2. Start RepoQL, then your app. Use the app, and check that telemetry is arriving.
  3. Find what failed. Group the errors, follow the request across services, and read the code.
  4. Fix it and prove it. The next run has no errors.
  5. Find what’s slow and make it fast. Rank the operations, read one request, and compare runs.

The example is a small shop with three ASP.NET Core services. storefront quotes a basket: for each item it asks catalog for the product, then pricing for its price. The services use Aspire’s standard service defaults, which already export OpenTelemetry.

Connect the AppHost

You need Aspire 13.4 or later, and rql installed (see Get started). This walkthrough used Aspire 13.5.4.

Add the RepoQL.Aspire package to your AppHost project.

Terminal

dotnet add Shopfront.AppHost package RepoQL.Aspire

Output

info : PackageReference for package 'RepoQL.Aspire' version '0.1.1' added to file '…/Shopfront.AppHost/Shopfront.AppHost.csproj'.

Then call AddRepoQL() in the AppHost.

Shopfront.AppHost/AppHost.cs

var builder = DistributedApplication.CreateBuilder(args);

builder.AddRepoQL();

var catalog = builder.AddProject<Projects.Catalog>("catalog");
var pricing = builder.AddProject<Projects.Pricing>("pricing");

builder.AddProject<Projects.Storefront>("storefront")
    .WithReference(catalog)
    .WithReference(pricing);

builder.Build().Run();

When the app starts, AddRepoQL() points every service’s OpenTelemetry exporter at RepoQL. RepoQL records everything and forwards it on to the Aspire dashboard, so the dashboard works as before. This only happens when you run the app. aspire publish output doesn’t change.

RepoQL forwards over HTTP, but the dashboard’s default telemetry endpoint speaks gRPC. So give the dashboard an HTTP endpoint in the launch profile you run. Any free port will do. In this walkthrough, dotnet run started the https profile.

Shopfront.AppHost/Properties/launchSettings.json

"environmentVariables": {
  …
  "ASPIRE_DASHBOARD_OTLP_ENDPOINT_URL": "https://localhost:21270",
  "ASPIRE_RESOURCE_SERVICE_ENDPOINT_URL": "https://localhost:22223",
  "ASPIRE_DASHBOARD_OTLP_HTTP_ENDPOINT_URL": "https://localhost:21271"
}

Without that last line, RepoQL still records everything, but the dashboard’s telemetry pages stay empty.

Start RepoQL, then your app

RepoQL.Aspire sends telemetry to the RepoQL host for your repository, so start the host first. Run this in the repository and leave it running.

Terminal

rql serve

If your agent is already using RepoQL in this repository, a host is running. Start your own anyway for a long session. A host started for an agent shuts down after about 15 minutes without a connected client, even while your app is still sending telemetry.

Now start the app as you usually do.

Terminal

dotnet run --project Shopfront.AppHost

Output

     AppHost:  Shopfront.AppHost.csproj
   Dashboard:  https://localhost:17031/login?t=…

In the Aspire dashboard, a repoql resource now runs beside your services.

Use the app. Here a customer asks for a quote with a lowercase product code. Ten six-item baskets follow.

Terminal

curl -s -o /dev/null -w 'HTTP %{http_code} in %{time_total}s\n' "http://localhost:5100/quote?skus=tea-104,MUG-220"
for i in $(seq 10); do curl -s -o /dev/null -w '%{http_code} %{time_total}\n' "http://localhost:5100/quote?skus=TEA-104,MUG-220,CUP-310,POT-415,TIN-502,SPN-618"; done

Output

HTTP 500 in 6.582251s
200 0.333600
200 0.319830
200 0.321171
200 0.320443
200 0.320364
200 0.321769
200 0.318232
200 0.315607
200 0.317558
200 0.319780

The lowercase quote failed after six and a half seconds. The rest took about a third of a second each. Check what RepoQL recorded.

Terminal

rql query "SELECT run_id, status, spans, errors FROM watch.summary()"

Output

run_id	status	spans	errors
58155afab32842d8ae9dbed8a394923d	live · last event 0s ago	261 spans, 11 traces, 96.6% ok, slowest GET /quote 6581.4ms, top GET /quote 9780.5ms, GET 3362.9ms, GET /products/{sku} 1911.4ms, 07:37:14-07:37:24	6 groups, top KeyNotFoundException @ System.Collections.Generic.Dictionary`2.get_Item ×4

Each start of the app is a new run. This one is live, with 11 traces for the 11 requests, so telemetry is arriving. The errors column already points at the failure.

Find what failed

Start with the errors, grouped.

Terminal

rql query "SELECT occurrences, services, error_type, message FROM watch.errors() ORDER BY first_seen"

Output

occurrences	services	error_type	message
4	pricing	KeyNotFoundException	The given key 'tea-104' was not present in the dictionary.
4	storefront	span-error	GET
4	pricing	span-error	GET /prices/{sku}
1	storefront	Error	Execution attempt. Source: '-standard//Standard-Retry', Operation Key: '', Result: '500', Handled: 'True', Attempt: '3', Execution Time: 22.3118ms
1	storefront	HttpRequestException	Response status code does not indicate success: 500 (Internal Server Error).
1	storefront	span-error	GET /quote

One failed request produced six groups in two services. The first is the cause: pricing couldn’t find the key tea-104, four times over. The rest are what followed in storefront: its retries, the 500 it received, and its own failed request.

Every group carries the trace_id of a request it happened in. Use it to see the whole request.

Terminal

rql query "SELECT trace_id FROM watch.errors() WHERE error_type = 'KeyNotFoundException'"

Output

trace_id
8e2b7ab4a01028e6fb4a0818db6f6ad2

Terminal

rql query "SELECT t.rendered, s.service_name AS service FROM watch.trace_tree('8e2b7ab4a01028e6fb4a0818db6f6ad2') t JOIN watch.otel_span s USING (span_pk) ORDER BY t.sort_path"

Output

rendered	service
GET /quote 6581.389ms ERROR	storefront
  GET 71.138ms	storefront
    GET /products/{sku} 55.776ms	catalog
  GET 59.946ms ERROR	storefront
    GET /prices/{sku} 47.975ms ERROR	pricing
  GET 21.899ms ERROR	storefront
    GET /prices/{sku} 21.048ms ERROR	pricing
  GET 21.281ms ERROR	storefront
    GET /prices/{sku} 20.632ms ERROR	pricing
  GET 22.084ms ERROR	storefront
    GET /prices/{sku} 21.504ms ERROR	pricing

catalog answered. Then pricing failed four times: the first call, plus three retries from Aspire’s standard resilience handler. Each attempt took about 20 ms, so nearly all of the 6.6 seconds went on waiting between retries. Retrying doesn’t help a failure that happens every time.

Now find the code. watch.error_frames() splits each stack trace into frames.

Terminal

rql query "SELECT e.services, f.file, f.line FROM watch.errors() e JOIN watch.error_frames() f USING (fingerprint) WHERE f.file IS NOT NULL ORDER BY e.first_seen"

Output

services	file	line
pricing	…/shopfront/Pricing/Program.cs	21
storefront	…/shopfront/Storefront/Program.cs	20

The paths are where the code was built, shortened here with …. Read both lines by their address in the repository.

Terminal

rql read "file:///Pricing/Program.cs#line=18,22"

Output

app.MapGet("/prices/{sku}", async (string sku) =>
{
    await Task.Delay(20); // the price book
    return Results.Ok(new Price(sku, prices[sku]));
});

Terminal

rql read "file:///Storefront/Program.cs#line=16,22"

Output

    var lines = new List<QuoteLine>();
    foreach (var sku in skus.Split(','))
    {
        var product = await catalog.GetFromJsonAsync<Product>($"/products/{sku}");
        var price = await pricing.GetFromJsonAsync<Price>($"/prices/{sku}");
        lines.Add(new QuoteLine(product!.Sku, product.Name, price!.Amount));
    }

The exception is in pricing, but pricing is right. Its price book is keyed by canonical product codes, like TEA-104. The catalog found tea-104 anyway, because it ignores case, and it returns the canonical code. The bug is in storefront: it prices the customer’s sku instead of the catalog’s product.Sku.

Fix it and prove it

Price the product the catalog returned.

Storefront/Program.cs

var price = await pricing.GetFromJsonAsync<Price>($"/prices/{product!.Sku}");

Stop the app and start it again. That begins a new run. Then send the same requests. The lowercase quote now works.

Terminal

curl -s -w '\nHTTP %{http_code} in %{time_total}s\n' "http://localhost:5100/quote?skus=tea-104,MUG-220"

Output

{"lines":[{"sku":"TEA-104","name":"Kettle, 1.7 litre","amount":34.00},{"sku":"MUG-220","name":"Stoneware mug","amount":12.50}],"total":46.50}
HTTP 200 in 0.362501s

Compare the two runs.

Terminal

rql query "SELECT run_id, status, spans, errors FROM watch.summary()"

Output

run_id	status	spans	errors
47c76a22767d4670b5b452873ed45f96	live · last event 0s ago	259 spans, 11 traces, 100.0% ok, slowest GET /quote 500.3ms, top GET /quote 4578.9ms, GET 4463.7ms, GET /products/{sku} 2133.5ms, 07:38:11-07:38:16	none
58155afab32842d8ae9dbed8a394923d	exited 0 · last event 32s ago	261 spans, 11 traces, 96.6% ok, slowest GET /quote 6581.4ms, top GET /quote 9780.5ms, GET 3362.9ms, GET /products/{sku} 1911.4ms, 07:37:14-07:37:24	6 groups, top KeyNotFoundException @ System.Collections.Generic.Dictionary`2.get_Item ×4

The new run is 100% ok, with no errors. The old run stays beside it, marked exited 0 when the app stopped. So the before and the after are both on record.

Find what’s slow and make it fast

With the failure gone, rank the new run’s operations.

Terminal

rql query "SELECT name, kind, count, round(p50_ms) AS p50_ms, round(p95_ms) AS p95_ms, round(total_ms) AS total_ms FROM watch.span_stats('47c76a22767d4670b5b452873ed45f96')"

Output

name	kind	count	p50_ms	p95_ms	total_ms
GET /quote	server	11	406	491	4579
GET	client	124	32	52	4464
GET /products/{sku}	server	62	32	40	2134
GET /prices/{sku}	server	62	23	30	1582

A quote takes 406 ms at the median. The calls it makes are quick: 32 ms to catalog and 23 ms to pricing. But 11 quotes made 124 calls. So look at how one quote arranges them.

Pick a typical quote, the median rather than the slowest. The slowest includes the app warming up.

Terminal

rql query "SELECT trace_id FROM watch.otel_span WHERE run_id = '47c76a22767d4670b5b452873ed45f96' AND name = 'GET /quote' ORDER BY duration_ms LIMIT 1 OFFSET 5"

Output

trace_id
12e26be896649f53ef272d5426235fb9

List its calls in the order they started.

Terminal

rql query "SELECT s.service_name AS service, s.name AS operation, round(s.duration_ms) AS ms, round((s.start_time_unix_nano - min(s.start_time_unix_nano) OVER ()) / 1e6) AS starts_at_ms FROM watch.trace_tree('12e26be896649f53ef272d5426235fb9') t JOIN watch.otel_span s USING (span_pk) WHERE t.depth <> 1 ORDER BY s.start_time_unix_nano"

Output

service	operation	ms	starts_at_ms
storefront	GET /quote	406	0
catalog	GET /products/{sku}	41	6
pricing	GET /prices/{sku}	24	67
catalog	GET /products/{sku}	31	91
pricing	GET /prices/{sku}	23	123
catalog	GET /products/{sku}	39	167
pricing	GET /prices/{sku}	21	216
catalog	GET /products/{sku}	30	246
pricing	GET /prices/{sku}	21	277
catalog	GET /products/{sku}	31	299
pricing	GET /prices/{sku}	21	331
catalog	GET /products/{sku}	30	353
pricing	GET /prices/{sku}	21	384

The calls run one after another: catalog, pricing, catalog, pricing, six times. Each waits for the one before, so the quote takes as long as all twelve added up. But the items don’t depend on each other, so look them all up at once.

Storefront/Program.cs

// Items are independent: look them all up at once.
var lines = await Task.WhenAll(skus.Split(',').Select(async sku =>
{
    var product = await catalog.GetFromJsonAsync<Product>($"/products/{sku}");
    var price = await pricing.GetFromJsonAsync<Price>($"/prices/{product!.Sku}");
    return new QuoteLine(product.Sku, product.Name, price!.Amount);
}));

Restart the app and send the same requests. Then list a typical quote’s calls from the new run.

Terminal

rql query "SELECT s.service_name AS service, s.name AS operation, round(s.duration_ms) AS ms, round((s.start_time_unix_nano - min(s.start_time_unix_nano) OVER ()) / 1e6) AS starts_at_ms FROM watch.trace_tree('bea7201106a4578917a39f9699d38ed5') t JOIN watch.otel_span s USING (span_pk) WHERE t.depth <> 1 ORDER BY s.start_time_unix_nano"

Output

service	operation	ms	starts_at_ms
storefront	GET /quote	65	0
catalog	GET /products/{sku}	31	1
catalog	GET /products/{sku}	30	1
catalog	GET /products/{sku}	30	1
catalog	GET /products/{sku}	30	1
catalog	GET /products/{sku}	30	1
catalog	GET /products/{sku}	30	11
pricing	GET /prices/{sku}	24	40
pricing	GET /prices/{sku}	24	40
pricing	GET /prices/{sku}	24	40
pricing	GET /prices/{sku}	24	40
pricing	GET /prices/{sku}	24	40
pricing	GET /prices/{sku}	20	44

Now the six catalog calls start together, then the six pricing calls. The quote takes 65 ms. Finally, compare every run in one query.

Terminal

rql query "SELECT run_id, count(*) AS quotes, count(*) FILTER (WHERE status_code = 'error') AS failed, round(median(duration_ms)) AS p50_ms, round(quantile_cont(duration_ms, 0.95)) AS p95_ms FROM watch.otel_span WHERE name = 'GET /quote' GROUP BY run_id ORDER BY min(start_time_unix_nano)"

Output

run_id	quotes	failed	p50_ms	p95_ms
58155afab32842d8ae9dbed8a394923d	11	1	319	3457
47c76a22767d4670b5b452873ed45f96	11	0	406	491
0ee1d12ccaad4b61bc477e1688a0dd2e	11	0	65	181

The median quote fell from 406 ms to 65 ms. The first run’s p95 is its failed request. Note the first two runs: the quotes did the same work, yet their medians differ by almost 90 ms, because timings on a laptop wander. A sixfold change is well clear of that. Repeat both runs before you claim a smaller gain.

Ask your agent

With RepoQL connected, your agent runs these same queries through its query and read tools. So you can ask in plain words:

RepoQL keeps a run for six hours after the app stops, so ask while it’s fresh.

What AddRepoQL() changes

At startup it registers a run with the RepoQL host, then sets these variables on every resource that exports OpenTelemetry.

VariableNew value
OTEL_EXPORTER_OTLP_ENDPOINTThe RepoQL host’s collector
OTEL_EXPORTER_OTLP_PROTOCOLhttp/protobuf
OTEL_EXPORTER_OTLP_HEADERSA header that routes the telemetry to this run
OTEL_RESOURCE_ATTRIBUTESAspire’s attributes, plus repoql.watch.run_id

The host records every payload before it forwards it. If the dashboard can’t take a payload, recording carries on and the failure is counted.

Terminal

rql query "SELECT run_id, target_url, forwarded, dropped, failed FROM watch.forward_stats"

Output

run_id	target_url	forwarded	dropped	failed
58155afab32842d8ae9dbed8a394923d	https://localhost:21271/	178	0	0
47c76a22767d4670b5b452873ed45f96	https://localhost:21271/	146	0	0
0ee1d12ccaad4b61bc477e1688a0dd2e	https://localhost:21271/	165	0	0

If RepoQL can’t be reached when the app starts, the repoql resource reports it, and your app starts with Aspire’s usual telemetry.

When it goes wrong

SymptomLikely causeWhat to do
The repoql resource shows Failed to start, and the AppHost’s log says “No host is running for this repository”No RepoQL host was running when the app startedRun rql serve in the repository, then restart the app
The repoql resource shows Failed to start, and the log says the rql CLI was not foundrql isn’t on the AppHost’s PATH or in ~/.local/binInstall RepoQL, or set REPOQL_CLI_PATH to the rql binary
RepoQL has the telemetry, but the dashboard’s telemetry pages are emptyThe profile you ran has no ASPIRE_DASHBOARD_OTLP_HTTP_ENDPOINT_URLAdd it to that profile, then check watch.forward_stats
Telemetry stops arriving partway through a sessionThe host was started for an agent, and shut down after about 15 minutes without a clientStart the host yourself with rql serve. It runs until you stop it
One service sends nothingIt doesn’t export OpenTelemetry, or its resource has no OTLP exporterUse Aspire’s service defaults, or add WithOtlpExporter() to the resource
A run has disappearedRuns are removed six hours after the app stopsSave the numbers you need while the run exists

To open RepoQL’s own dashboard, run rql dashboard in the repository.

What you’ve done

You connected an Aspire app to RepoQL with one line. Then you worked the way your agent can: from grouped errors, to the request that crossed services, to the line that caused it. You proved the fix with a clean run, and the speed-up with one query across runs.

Further reference

The watch tool page and the Profile with watch guide. The package is RepoQL.Aspire on GitHub and on NuGet. In the installed manual, help:///tools/watch/watch.md covers rql watch env, which the package uses.