Integration
Observe, fix and optimise with Aspire
Send every service’s telemetry from your Aspire app to RepoQL. Your agent can then find what failed and what’s slow across the whole system, follow each one to the line of code, and prove a fix with a before-and-after query.
The route
- Connect the AppHost. Add one package and one line.
- Start RepoQL, then your app. Use the app, and check that telemetry is arriving.
- Find what failed. Group the errors, follow the request across services, and read the code.
- Fix it and prove it. The next run has no errors.
- Find what’s slow and make it fast. Rank the operations, read one request, and compare runs.
The example is a small shop with three ASP.NET Core services. storefront quotes a basket: for each item it asks catalog for the product, then pricing for its price. The services use Aspire’s standard service defaults, which already export OpenTelemetry.
Connect the AppHost
You need Aspire 13.4 or later, and rql installed (see Get started). This walkthrough used Aspire 13.5.4.
Add the RepoQL.Aspire package to your AppHost project.
Terminal
dotnet add Shopfront.AppHost package RepoQL.Aspire
Output
info : PackageReference for package 'RepoQL.Aspire' version '0.1.1' added to file '…/Shopfront.AppHost/Shopfront.AppHost.csproj'.
Then call AddRepoQL() in the AppHost.
Shopfront.AppHost/AppHost.cs
var builder = DistributedApplication.CreateBuilder(args);
builder.AddRepoQL();
var catalog = builder.AddProject<Projects.Catalog>("catalog");
var pricing = builder.AddProject<Projects.Pricing>("pricing");
builder.AddProject<Projects.Storefront>("storefront")
.WithReference(catalog)
.WithReference(pricing);
builder.Build().Run();
When the app starts, AddRepoQL() points every service’s OpenTelemetry exporter at RepoQL. RepoQL records everything and forwards it on to the Aspire dashboard, so the dashboard works as before. This only happens when you run the app. aspire publish output doesn’t change.
RepoQL forwards over HTTP, but the dashboard’s default telemetry endpoint speaks gRPC. So give the dashboard an HTTP endpoint in the launch profile you run. Any free port will do. In this walkthrough, dotnet run started the https profile.
Shopfront.AppHost/Properties/launchSettings.json
"environmentVariables": {
…
"ASPIRE_DASHBOARD_OTLP_ENDPOINT_URL": "https://localhost:21270",
"ASPIRE_RESOURCE_SERVICE_ENDPOINT_URL": "https://localhost:22223",
"ASPIRE_DASHBOARD_OTLP_HTTP_ENDPOINT_URL": "https://localhost:21271"
}
Without that last line, RepoQL still records everything, but the dashboard’s telemetry pages stay empty.
Start RepoQL, then your app
RepoQL.Aspire sends telemetry to the RepoQL host for your repository, so start the host first. Run this in the repository and leave it running.
Terminal
rql serve
If your agent is already using RepoQL in this repository, a host is running. Start your own anyway for a long session. A host started for an agent shuts down after about 15 minutes without a connected client, even while your app is still sending telemetry.
Now start the app as you usually do.
Terminal
dotnet run --project Shopfront.AppHost
Output
AppHost: Shopfront.AppHost.csproj
Dashboard: https://localhost:17031/login?t=…
In the Aspire dashboard, a repoql resource now runs beside your services.
Use the app. Here a customer asks for a quote with a lowercase product code. Ten six-item baskets follow.
Terminal
curl -s -o /dev/null -w 'HTTP %{http_code} in %{time_total}s\n' "http://localhost:5100/quote?skus=tea-104,MUG-220"
for i in $(seq 10); do curl -s -o /dev/null -w '%{http_code} %{time_total}\n' "http://localhost:5100/quote?skus=TEA-104,MUG-220,CUP-310,POT-415,TIN-502,SPN-618"; done
Output
HTTP 500 in 6.582251s
200 0.333600
200 0.319830
200 0.321171
200 0.320443
200 0.320364
200 0.321769
200 0.318232
200 0.315607
200 0.317558
200 0.319780
The lowercase quote failed after six and a half seconds. The rest took about a third of a second each. Check what RepoQL recorded.
Terminal
rql query "SELECT run_id, status, spans, errors FROM watch.summary()"
Output
run_id status spans errors
58155afab32842d8ae9dbed8a394923d live · last event 0s ago 261 spans, 11 traces, 96.6% ok, slowest GET /quote 6581.4ms, top GET /quote 9780.5ms, GET 3362.9ms, GET /products/{sku} 1911.4ms, 07:37:14-07:37:24 6 groups, top KeyNotFoundException @ System.Collections.Generic.Dictionary`2.get_Item ×4
Each start of the app is a new run. This one is live, with 11 traces for the 11 requests, so telemetry is arriving. The errors column already points at the failure.
Find what failed
Start with the errors, grouped.
Terminal
rql query "SELECT occurrences, services, error_type, message FROM watch.errors() ORDER BY first_seen"
Output
occurrences services error_type message
4 pricing KeyNotFoundException The given key 'tea-104' was not present in the dictionary.
4 storefront span-error GET
4 pricing span-error GET /prices/{sku}
1 storefront Error Execution attempt. Source: '-standard//Standard-Retry', Operation Key: '', Result: '500', Handled: 'True', Attempt: '3', Execution Time: 22.3118ms
1 storefront HttpRequestException Response status code does not indicate success: 500 (Internal Server Error).
1 storefront span-error GET /quote
One failed request produced six groups in two services. The first is the cause: pricing couldn’t find the key tea-104, four times over. The rest are what followed in storefront: its retries, the 500 it received, and its own failed request.
Every group carries the trace_id of a request it happened in. Use it to see the whole request.
Terminal
rql query "SELECT trace_id FROM watch.errors() WHERE error_type = 'KeyNotFoundException'"
Output
trace_id
8e2b7ab4a01028e6fb4a0818db6f6ad2
Terminal
rql query "SELECT t.rendered, s.service_name AS service FROM watch.trace_tree('8e2b7ab4a01028e6fb4a0818db6f6ad2') t JOIN watch.otel_span s USING (span_pk) ORDER BY t.sort_path"
Output
rendered service
GET /quote 6581.389ms ERROR storefront
GET 71.138ms storefront
GET /products/{sku} 55.776ms catalog
GET 59.946ms ERROR storefront
GET /prices/{sku} 47.975ms ERROR pricing
GET 21.899ms ERROR storefront
GET /prices/{sku} 21.048ms ERROR pricing
GET 21.281ms ERROR storefront
GET /prices/{sku} 20.632ms ERROR pricing
GET 22.084ms ERROR storefront
GET /prices/{sku} 21.504ms ERROR pricing
catalog answered. Then pricing failed four times: the first call, plus three retries from Aspire’s standard resilience handler. Each attempt took about 20 ms, so nearly all of the 6.6 seconds went on waiting between retries. Retrying doesn’t help a failure that happens every time.
Now find the code. watch.error_frames() splits each stack trace into frames.
Terminal
rql query "SELECT e.services, f.file, f.line FROM watch.errors() e JOIN watch.error_frames() f USING (fingerprint) WHERE f.file IS NOT NULL ORDER BY e.first_seen"
Output
services file line
pricing …/shopfront/Pricing/Program.cs 21
storefront …/shopfront/Storefront/Program.cs 20
The paths are where the code was built, shortened here with …. Read both lines by their address in the repository.
Terminal
rql read "file:///Pricing/Program.cs#line=18,22"
Output
app.MapGet("/prices/{sku}", async (string sku) =>
{
await Task.Delay(20); // the price book
return Results.Ok(new Price(sku, prices[sku]));
});
Terminal
rql read "file:///Storefront/Program.cs#line=16,22"
Output
var lines = new List<QuoteLine>();
foreach (var sku in skus.Split(','))
{
var product = await catalog.GetFromJsonAsync<Product>($"/products/{sku}");
var price = await pricing.GetFromJsonAsync<Price>($"/prices/{sku}");
lines.Add(new QuoteLine(product!.Sku, product.Name, price!.Amount));
}
The exception is in pricing, but pricing is right. Its price book is keyed by canonical product codes, like TEA-104. The catalog found tea-104 anyway, because it ignores case, and it returns the canonical code. The bug is in storefront: it prices the customer’s sku instead of the catalog’s product.Sku.
Fix it and prove it
Price the product the catalog returned.
Storefront/Program.cs
var price = await pricing.GetFromJsonAsync<Price>($"/prices/{product!.Sku}");
Stop the app and start it again. That begins a new run. Then send the same requests. The lowercase quote now works.
Terminal
curl -s -w '\nHTTP %{http_code} in %{time_total}s\n' "http://localhost:5100/quote?skus=tea-104,MUG-220"
Output
{"lines":[{"sku":"TEA-104","name":"Kettle, 1.7 litre","amount":34.00},{"sku":"MUG-220","name":"Stoneware mug","amount":12.50}],"total":46.50}
HTTP 200 in 0.362501s
Compare the two runs.
Terminal
rql query "SELECT run_id, status, spans, errors FROM watch.summary()"
Output
run_id status spans errors
47c76a22767d4670b5b452873ed45f96 live · last event 0s ago 259 spans, 11 traces, 100.0% ok, slowest GET /quote 500.3ms, top GET /quote 4578.9ms, GET 4463.7ms, GET /products/{sku} 2133.5ms, 07:38:11-07:38:16 none
58155afab32842d8ae9dbed8a394923d exited 0 · last event 32s ago 261 spans, 11 traces, 96.6% ok, slowest GET /quote 6581.4ms, top GET /quote 9780.5ms, GET 3362.9ms, GET /products/{sku} 1911.4ms, 07:37:14-07:37:24 6 groups, top KeyNotFoundException @ System.Collections.Generic.Dictionary`2.get_Item ×4
The new run is 100% ok, with no errors. The old run stays beside it, marked exited 0 when the app stopped. So the before and the after are both on record.
Find what’s slow and make it fast
With the failure gone, rank the new run’s operations.
Terminal
rql query "SELECT name, kind, count, round(p50_ms) AS p50_ms, round(p95_ms) AS p95_ms, round(total_ms) AS total_ms FROM watch.span_stats('47c76a22767d4670b5b452873ed45f96')"
Output
name kind count p50_ms p95_ms total_ms
GET /quote server 11 406 491 4579
GET client 124 32 52 4464
GET /products/{sku} server 62 32 40 2134
GET /prices/{sku} server 62 23 30 1582
A quote takes 406 ms at the median. The calls it makes are quick: 32 ms to catalog and 23 ms to pricing. But 11 quotes made 124 calls. So look at how one quote arranges them.
Pick a typical quote, the median rather than the slowest. The slowest includes the app warming up.
Terminal
rql query "SELECT trace_id FROM watch.otel_span WHERE run_id = '47c76a22767d4670b5b452873ed45f96' AND name = 'GET /quote' ORDER BY duration_ms LIMIT 1 OFFSET 5"
Output
trace_id
12e26be896649f53ef272d5426235fb9
List its calls in the order they started.
Terminal
rql query "SELECT s.service_name AS service, s.name AS operation, round(s.duration_ms) AS ms, round((s.start_time_unix_nano - min(s.start_time_unix_nano) OVER ()) / 1e6) AS starts_at_ms FROM watch.trace_tree('12e26be896649f53ef272d5426235fb9') t JOIN watch.otel_span s USING (span_pk) WHERE t.depth <> 1 ORDER BY s.start_time_unix_nano"
Output
service operation ms starts_at_ms
storefront GET /quote 406 0
catalog GET /products/{sku} 41 6
pricing GET /prices/{sku} 24 67
catalog GET /products/{sku} 31 91
pricing GET /prices/{sku} 23 123
catalog GET /products/{sku} 39 167
pricing GET /prices/{sku} 21 216
catalog GET /products/{sku} 30 246
pricing GET /prices/{sku} 21 277
catalog GET /products/{sku} 31 299
pricing GET /prices/{sku} 21 331
catalog GET /products/{sku} 30 353
pricing GET /prices/{sku} 21 384
The calls run one after another: catalog, pricing, catalog, pricing, six times. Each waits for the one before, so the quote takes as long as all twelve added up. But the items don’t depend on each other, so look them all up at once.
Storefront/Program.cs
// Items are independent: look them all up at once.
var lines = await Task.WhenAll(skus.Split(',').Select(async sku =>
{
var product = await catalog.GetFromJsonAsync<Product>($"/products/{sku}");
var price = await pricing.GetFromJsonAsync<Price>($"/prices/{product!.Sku}");
return new QuoteLine(product.Sku, product.Name, price!.Amount);
}));
Restart the app and send the same requests. Then list a typical quote’s calls from the new run.
Terminal
rql query "SELECT s.service_name AS service, s.name AS operation, round(s.duration_ms) AS ms, round((s.start_time_unix_nano - min(s.start_time_unix_nano) OVER ()) / 1e6) AS starts_at_ms FROM watch.trace_tree('bea7201106a4578917a39f9699d38ed5') t JOIN watch.otel_span s USING (span_pk) WHERE t.depth <> 1 ORDER BY s.start_time_unix_nano"
Output
service operation ms starts_at_ms
storefront GET /quote 65 0
catalog GET /products/{sku} 31 1
catalog GET /products/{sku} 30 1
catalog GET /products/{sku} 30 1
catalog GET /products/{sku} 30 1
catalog GET /products/{sku} 30 1
catalog GET /products/{sku} 30 11
pricing GET /prices/{sku} 24 40
pricing GET /prices/{sku} 24 40
pricing GET /prices/{sku} 24 40
pricing GET /prices/{sku} 24 40
pricing GET /prices/{sku} 24 40
pricing GET /prices/{sku} 20 44
Now the six catalog calls start together, then the six pricing calls. The quote takes 65 ms. Finally, compare every run in one query.
Terminal
rql query "SELECT run_id, count(*) AS quotes, count(*) FILTER (WHERE status_code = 'error') AS failed, round(median(duration_ms)) AS p50_ms, round(quantile_cont(duration_ms, 0.95)) AS p95_ms FROM watch.otel_span WHERE name = 'GET /quote' GROUP BY run_id ORDER BY min(start_time_unix_nano)"
Output
run_id quotes failed p50_ms p95_ms
58155afab32842d8ae9dbed8a394923d 11 1 319 3457
47c76a22767d4670b5b452873ed45f96 11 0 406 491
0ee1d12ccaad4b61bc477e1688a0dd2e 11 0 65 181
The median quote fell from 406 ms to 65 ms. The first run’s p95 is its failed request. Note the first two runs: the quotes did the same work, yet their medians differ by almost 90 ms, because timings on a laptop wander. A sixfold change is well clear of that. Repeat both runs before you claim a smaller gain.
Ask your agent
With RepoQL connected, your agent runs these same queries through its query and read tools. So you can ask in plain words:
- “Use RepoQL’s watch telemetry for the latest run of this app. What failed, which service caused it, and which line of code?”
- “Which operation is slowest in the latest run, and why? Show me one typical request.”
- “Compare
GET /quoteacross the last two runs.”
RepoQL keeps a run for six hours after the app stops, so ask while it’s fresh.
What AddRepoQL() changes
At startup it registers a run with the RepoQL host, then sets these variables on every resource that exports OpenTelemetry.
| Variable | New value |
|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT | The RepoQL host’s collector |
OTEL_EXPORTER_OTLP_PROTOCOL | http/protobuf |
OTEL_EXPORTER_OTLP_HEADERS | A header that routes the telemetry to this run |
OTEL_RESOURCE_ATTRIBUTES | Aspire’s attributes, plus repoql.watch.run_id |
The host records every payload before it forwards it. If the dashboard can’t take a payload, recording carries on and the failure is counted.
Terminal
rql query "SELECT run_id, target_url, forwarded, dropped, failed FROM watch.forward_stats"
Output
run_id target_url forwarded dropped failed
58155afab32842d8ae9dbed8a394923d https://localhost:21271/ 178 0 0
47c76a22767d4670b5b452873ed45f96 https://localhost:21271/ 146 0 0
0ee1d12ccaad4b61bc477e1688a0dd2e https://localhost:21271/ 165 0 0
If RepoQL can’t be reached when the app starts, the repoql resource reports it, and your app starts with Aspire’s usual telemetry.
When it goes wrong
| Symptom | Likely cause | What to do |
|---|---|---|
The repoql resource shows Failed to start, and the AppHost’s log says “No host is running for this repository” | No RepoQL host was running when the app started | Run rql serve in the repository, then restart the app |
The repoql resource shows Failed to start, and the log says the rql CLI was not found | rql isn’t on the AppHost’s PATH or in ~/.local/bin | Install RepoQL, or set REPOQL_CLI_PATH to the rql binary |
| RepoQL has the telemetry, but the dashboard’s telemetry pages are empty | The profile you ran has no ASPIRE_DASHBOARD_OTLP_HTTP_ENDPOINT_URL | Add it to that profile, then check watch.forward_stats |
| Telemetry stops arriving partway through a session | The host was started for an agent, and shut down after about 15 minutes without a client | Start the host yourself with rql serve. It runs until you stop it |
| One service sends nothing | It doesn’t export OpenTelemetry, or its resource has no OTLP exporter | Use Aspire’s service defaults, or add WithOtlpExporter() to the resource |
| A run has disappeared | Runs are removed six hours after the app stops | Save the numbers you need while the run exists |
To open RepoQL’s own dashboard, run rql dashboard in the repository.
What you’ve done
You connected an Aspire app to RepoQL with one line. Then you worked the way your agent can: from grouped errors, to the request that crossed services, to the line that caused it. You proved the fix with a clean run, and the speed-up with one query across runs.
Further reference
The watch tool page and the Profile with watch guide. The package is RepoQL.Aspire on GitHub and on NuGet. In the installed manual, help:///tools/watch/watch.md covers rql watch env, which the package uses.