NorthLedger InsightsRashadul Islam Roman Bookaudit

City data to 2026-09-21 (latest evaluation in the file) page rebuilt 2026-10-08 16:11:51 UTC, by hand, no schedule 47 build checks passed

Data analytics and automation, Toronto

Messy data in.
Checked decisions out.

I turn scattered business data into cleaned tables, checked findings and honest forecasts. This page runs the whole method on public City of Toronto data about rental apartment buildings: every figure about the buildings is computed by code from the City's data, and every claim links to its receipt.

Research collaborationFor researchers, labs and data teams.

Try it on your own file Drop a messy CSV and get a checked report like this one, made inside your browser.

The report

RentSafeTO Building Health, as an interactive report

A Power BI-style report built in the browser from City of Toronto open data, with evaluations to 2026-09-21. Use the slicers: the cards, the Portfolio pulse and Pillars pages and the scorecard below filter together; Points per fix and the risk chart follow the district and ward only, and the trend chart follows a single ward. Tap or click a ward bar or map area to filter to it; tap another to switch. Clear a filter with the × on its chip or with Reset filters. With a keyboard, Ctrl or Cmd adds more wards and Esc clears the last filter. Public view is ward and district level only: no landlord, firm or address is ranked, and a view of fewer than five buildings shows no values.

Independent analysis by NorthLedger Insights of the City's open data. It is not produced by, affiliated with or endorsed by the City of Toronto or RentSafeTO.

RentSafeTO Building HealthCity of Toronto open data to 2026-09-21 (latest evaluation in the file)
Buildings
3,588
latest City evaluation of each
Units covered
327,317
confirmed units in those buildings
Average City score, unit-weighted
91.2
simple mean 90.4
Green sign
82.5%
yellow 16.5%, red 1.0%
Evaluations with a refused or blocked area, per hundred
3.99
an area the officer cannot enter scores zero
Audit zone, 2026 citywide, not filtered
37
of 1,380 evaluations that year at or below the cutoff score of 72 (proactive score only)
Evaluation due within ninety days citywide, not filtered
50
last evaluation plus two years; 413 more are already past that mark

Two cards are citywide and do not follow the slicers. The City prioritises the lowest 2.5% of evaluated buildings each year for audit and now uses proactive AND reactive scores; this count uses the proactive score only. Past the two-year mark means: buildings whose latest evaluation in this file is more than 2 years before the as-of date; the next evaluation may be scheduled or not yet published.

Average City score by ward, highest first
Ward map, shaded by average City score
Sign bands by property type

What this view says

Scorecard

Every ward, every pillar, one table

25 wards by 8 building-health pillars, then the City's own overall score, the share of buildings holding a green sign, refusals and rank. District and City rows sit underneath. Cells use the City's lobby-sign bands (source: the City's sign page) and always print the number and a band symbol, so colour is never the only signal. Inspired by the World Economic Forum's Travel and Tourism Development Index table layout; none of its data is reproduced.

Colour by
Zero means
What if a pillar counted more? Weight sliders

Ranks a weighted average of the eight pillar cells. The City's overall score is not changed; this shows how sensitive a ward's position is to what you care about.

Green 85 to 100Yellow 70 to 84Red 0 to 69
Toronto wards by building-health pillar: the average of each ward's buildings, latest City evaluation of each, with district and City rows. Each cell prints its score and its City sign band.
86.1 green90.4 green95.4 green95.5 green87.0 green91.2 green91.1 green89.8 green90.5 green80%0.012
89.5 green91.1 green97.7 green97.4 green90.0 green94.4 green94.7 green94.1 green93.2 green93%3.12
87.4 green91.2 green97.1 green99.0 green88.2 green94.6 green95.7 green91.2 green92.4 green91%4.04
87.7 green91.0 green94.6 green95.7 green86.5 green91.4 green92.7 green88.0 green90.5 green87%4.913
82.8 yellow87.8 green93.7 green93.6 green84.0 yellow85.7 green91.6 green88.0 green87.9 green70%1.322
87.9 green87.8 green95.1 green94.6 green88.5 green90.0 green93.0 green92.5 green90.7 green84%3.411
87.3 green85.7 green95.4 green93.4 green87.5 green90.3 green94.3 green91.9 green90.3 green81%0.014
88.5 green89.9 green96.4 green94.3 green89.1 green93.7 green96.3 green90.3 green92.0 green84%2.56
79.5 yellow84.7 yellow93.4 green90.3 green82.1 yellow84.3 yellow90.7 green88.0 green85.6 green60%9.825
89.0 green90.6 green97.7 green91.3 green88.0 green84.9 yellow91.4 green95.2 green89.9 green75%8.216
86.0 green88.2 green95.7 green91.2 green81.8 yellow85.7 green91.7 green92.6 green87.9 green76%11.623
87.3 green89.8 green94.9 green95.6 green85.6 green91.1 green92.6 green92.5 green90.7 green84%5.010
87.2 green89.7 green94.3 green94.0 green86.3 green84.4 yellow89.6 green95.8 green89.2 green76%6.119
87.7 green89.6 green93.5 green92.1 green88.6 green84.6 yellow89.7 green92.7 green89.2 green79%6.820
89.3 green93.8 green95.6 green96.1 green89.7 green92.1 green96.4 green95.0 green93.0 green95%0.63
84.6 yellow90.9 green95.5 green94.2 green91.6 green91.1 green95.1 green93.5 green91.4 green90%0.68
87.5 green91.2 green93.0 green92.2 green92.3 green93.4 green94.8 green90.9 green91.6 green85%2.77
84.3 yellow88.5 green95.2 green92.4 green92.4 green90.4 green94.2 green87.0 green90.1 green76%0.015
83.0 yellow92.0 green96.4 green96.8 green88.1 green92.2 green95.8 green92.4 green91.0 green88%8.99
82.2 yellow88.7 green94.2 green95.9 green86.4 green87.8 green95.3 green94.3 green89.7 green82%4.917
81.7 yellow88.5 green95.0 green94.5 green87.5 green88.7 green94.7 green91.9 green89.5 green80%0.018
88.0 green90.6 green96.6 green93.8 green90.9 green90.3 green96.8 green94.4 green92.0 green90%2.45
93.3 green90.2 green99.0 green98.3 green95.2 green90.5 green92.2 green87.4 green93.3 green89%0.01
86.5 green87.4 green91.1 green92.1 green88.4 green87.5 green89.8 green91.4 green88.6 green77%0.021
88.1 green86.9 green87.7 green89.5 green80.2 yellow79.8 yellow88.7 green88.7 green86.0 green57%0.024
Etobicoke Yorkaverage of its wards86.6 green89.2 green95.9 green95.8 green87.3 green91.2 green93.5 green91.0 green90.9 green83%1.7not ranked
North Yorkaverage of its wards87.0 green90.4 green95.1 green94.0 green90.6 green91.8 green95.0 green91.5 green91.5 green86%1.6not ranked
Toronto and East Yorkaverage of its wards85.9 green89.4 green95.1 green93.4 green85.9 green87.3 green91.8 green92.2 green89.2 green78%7.7not ranked
Scarboroughaverage of its wards86.6 green88.7 green93.9 green94.0 green88.1 green87.4 green92.9 green91.3 green89.9 green79%1.2not ranked
City of Torontoaverage of the 25 ward cells86.5 green89.4 green95.0 green94.2 green87.8 green89.2 green93.2 green91.6 green90.2 green81%3.5not ranked

The City row averages the 25 ward cells, each ward counted once, so it differs from the building-level figures in the report above: simple mean 90.4, unit-weighted 91.2, 82.5% in the green band and 3.99 evaluations with a refusal per hundred.

Ranked list, one pillar at a time
How each cell is built, and the limits
  • Each building's LATEST post-2023 evaluation (the RSN panel is collapsed to the latest row by completion date).
  • Overall = the City's CURRENT BUILDING EVAL SCORE (proactive score plus current reactive deductions), which is what sets the lobby sign colour. Not recomputed.
  • Pillar % for one building = sum(w_i x s_i / 3) / sum(w_i) x 100 over the pillar's items scored 1-3, with the City's published tier weights (High 3, Moderate 2, Cosmetic 0.5). N/A items are left out. Provisional choice, to be confirmed: items scored 0 (an area refused or blocked) are also left out and counted in the refusal columns instead; the *_mean_city_basis fields count them as 0 points, which is what the City's own score does (see analysis A).
  • Ward cell (default) = simple mean over the ward's buildings, each building counted once (*_mean). The unit-weighted mean (*_unit_weighted, weights = CONFIRMED UNITS) and the building count (*_n) are also given.
  • District and City rows: the default (*__avg_of_wards) is the SIMPLE AVERAGE OF THE MEMBER WARD CELLS, each ward counted once, like a regional-average row in the TTDI table. The pooled mean over all the buildings in the district or city (*__pooled_buildings) is also given; the two differ because wards hold different numbers of buildings.
  • Colour bands are the City's lobby-sign bands: 85-100 green, 70-84 yellow, 0-69 red.
  • pct_green/yellow/red = share of the ward's buildings whose latest City score is in the band.
  • refusal_evals_per_100 = latest evaluations with at least one item scored 0 per 100 evaluations; refused_items_per_100_evals = items scored 0 per 100 evaluations.
  • rank = 1 + number of wards with a higher overall_mean (ties share a rank); *_rank the same per pillar on *_mean.
  • low_n is true when a ward has fewer than 5 buildings.
  • The 8-pillar grouping is ours; the City publishes items and tiers, not pillars.
  • Scores are the City's bylaw-evaluation scores, not a safety rating; a red sign does not mean unsafe.
  • Latest evaluations span several years (2023-2026), so wards are compared across slightly different dates.
  • Provisional choice, to be confirmed: a refused area is kept as a separate flag here; the City's own score counts it as zero points.
  • Public view is ward/district level only; no landlord, firm or address is ranked.

Source: City of Toronto, Apartment Building Evaluation, via the CKAN portal; see receipts for the download time and checksum.

Run rental buildings? See what an audit of your own buildings covers, and the prices.

Analysis gallery

Four deep-dives, each with its limits

The kind of monthly deep-dive the retainer includes. Each card runs Question, Method, Chart, Finding, Limits and how to reproduce it. All of it is numpy and pandas code in this repository, written with AI assistance; it uses no model library and calls no AI model when it runs.

Analysis A

Points per fix: rebuilding the City's score

Question
Can the City's published proactive score be rebuilt from the 50 item scores and the published tier weights, and how does the City treat an item scored 0 (refused or obstructed)? Which fixes are worth the most points?
Method
Constrained: the published tier weights with two treatments of 0, compared with the published PROACTIVE BUILDING SCORE on every clean post-2023 row. Unconstrained: least squares on the 50 item deficits with a temporal train/test split, to see whether the data recover the tier ordering without being told it. score = round_half_up(round(sum(w_i x s_i / 3) / sum(w_i) x 100, 2)), over items that are not N/A; w = 3 / 2 / 0.5 by tier.
How a pillar score is built: each item's score out of three times its tier weight, summed, divided by the weights of the items that apply Item score1, 2 or 3 (0 = refused) Tier weightHigh 3, Moderate 2, Cosmetic 0.5 weight x score / 3summed over items divided by weightsof items that apply
N/A items leave both the top and the bottom of the fraction; a refused item stays in the bottom, which is why it costs full marks.
Expected City points per building from bringing an item to full marks, top items
Finding
The City's published weights reproduce 6,666 of 6,681 scores exactly (99.8%). On rows with a refused item, counting it as zero reproduces 99.4%; leaving it out reproduces 5.0%, so the City counts a refused area as zero points. A free least-squares fit tested on the latest evaluations scores R squared 0.980 with a mean error of 0.76 points, against 0.01 for the published formula, and gives a high-to-cosmetic ratio of 4.18 against the published 6. The same fit on scores produced by the City's exact formula also gives 4.18, so the gap is the straight-line fit's approximation, not a sign that the published weights are wrong. 15 evaluations stay unexplained.
Limits
  • The rebuild uses the tier weights published on the City's evaluations page; it cannot see officer notes or re-evaluations, which may explain the unexplained rows.
  • Points per fix assumes the other items stay the same and that an item can be raised to 3; it is not a cost model.
  • The unconstrained fit uses item deficits only; the City score is a ratio, so a linear fit is an approximation. The same fit on scores produced by the City's exact formula also gives a high-to-cosmetic ratio of 4.18, so the gap from the published 6 comes from that approximation; the data agree with the published weights. The fit is tested on the most recent 20% of evaluations (from 2026-06-23), not on each building's latest evaluation.
Reproduce
Download the 4 City files listed under Receipts into one folder, then run python build_rentsafe.py --check --input-dir <folder> (script). It rebuilds the scorecard, the cube, the trend, the ward map and Analyses A, B and D from those raw files in a temporary folder and runs their invariant checks; the data-health counts come from the engine-cleaned run and are not rebuilt by it. Their checksums are in the same table.

Analysis B

Will a building lose its green sign? A model against persistence

Question
Will a building's next biennial evaluation score fall below 85 (lose or miss the green sign)?
Method
numpy Newton-Raphson logistic regression, L2 lambda=100 (chosen on an inner temporal split of the training pairs at 2025-08-01); next-score model: numpy ridge regression, lambda=300. Pairs of consecutive evaluations of the same building: 3,093, split in time (temporal: pairs whose NEXT evaluation completed before 2026-01-01 train; on/after test), 1,721 to train and 1,372 to test. Baseline: P(next < 85) = training-set rate of next < 85 among pairs whose last score is in the same band [0, 70, 80, 85, 88, 91, 94, 101]. The model is champion only if it beats persistence on BOTH Brier (P(next<85)) and MAE (next score) out of sample.
Reliability on the test set: predicted chance against what happened, model and persistence
Test-set scores, model against simple baselines
Measure, test setModelBaselineBetter
Brier score for the next score falling below green (how far the predicted chances were from what happened; lower is better)0.12490.1211
band's past rate
baseline
Precision-recall area (how well it ranks the buildings that did fall; higher is better)0.3290.308
band's past rate
model
Mean absolute error of the next score, points (lower is better)4.977.11
repeat the last score
model
Mean absolute error of the next score, points (lower is better)4.975.08
last score plus the band's usual change
model
Finding
The model was not offered: it beat repeating the last score on average miss (4.97 against 7.11 points), but not the band's past rate on Brier score (0.125 against 0.121), and the pre-set rule needed both. A persistence baseline that adds the usual change for the building's band nearly matches the model on the score itself (5.08 points). The model's own range for the next score held 71.7% of the time (Wilson interval 69.3% to 74.0%) when it should hold about four in five. So the page offers no per-building prediction; for a rough expectation, the last score plus the usual change for its band is the practical guide.
Limits
  • About one post-2023 pair per building; the test set is one season (2026) and small.
  • Biennial evaluations: a building's condition between evaluations is not observed.
  • Registration attributes are a 2026 snapshot (possible look-ahead).
  • Reliability bins with few buildings are noisy; read them with their n.
Reproduce
The same command as Analysis A: python build_rentsafe.py --check --input-dir <folder>, with the 4 City files from Receipts in that folder.

Analysis C, the Forecast Lab

U.S. retail trade and food services: which forecast holds up?

Latest month, seasonally adjusted, Aug 2026
$773.9B
month on month +1.24%, year on year +6.01%
Winner out of sample
Seasonal naive with drift
a baseline; average miss 1.27% of the actual (MAPE); MASE 0.45, its miss as a share of the seasonal-naive miss. Picked on the same origins it is scored on: re-picked at each origin from only the errors known then, the forecasts missed by 1.51%, 0.24 points more
Band check, 80% band
82.3%
held 237 of 288 times (Wilson interval 77.5% to 86.3%; allowing for overlapping forecasts, 68.3% to 93.8%)
Next month, Sep 2026, not seasonally adjusted
$738.8B
Aug 2026 on the same basis $784.8B; step -5.9% (the usual seasonal move), +0.55% after seasonal adjustment; band $718.2B to $751.9B; backtest verdict RECOMMEND: champion chosen out of sample; its 80% band has not been shown to be off
Forecast from the latest month with its band, against history
How a rolling-origin backtest works: at each past month the method sees only the data up to that month, forecasts the next months, and is scored against what was later reported; then the origin moves forward one month history the method may seeforecast, scored later each row is one origin; the origin steps forward a month at a time, and no forecast sees its own future
Out-of-sample accuracy of every method, best first
MethodKindMAPE, percentMASEForecasts scored
Seasonal naive with drift championbaseline1.270.45288
Damped Holt (log, deseasonalised)fitted1.650.60288
Log-linear trend + seasonfitted1.810.66288
Seasonal naivebaseline3.481.28288
Driftbaseline6.002.15288
Naive (last value)baseline6.352.30288
Method
  • Data: FRED series RSAFSNA (U.S. retail trade and food services, not seasonally adjusted, millions of dollars), 416 monthly values from Jan 1992 to Aug 2026.
  • Pandemic months (Mar 2020 to Jun 2021) are treated as missing whenever a model estimates anything: averages skip changes that touch them, the log-linear fit leaves them out, and the Holt smoother carries its state across them and restarts its level on the first month after. No scored forecast starts from, looks back into, or lands in that window.
  • Six methods, all on the log of sales and all refit at every origin using only the data available at that point: naive, seasonal naive, drift, seasonal naive with drift, damped Holt and a log-linear trend plus month effects.
  • Rolling-origin backtest: 24 origins (Sep 2023 to Aug 2025), each forecasting 1 to 12 months ahead, so 288 forecasts per method, each compared with what was later reported. Methods are ranked by MASE (error relative to a seasonal-naive yardstick); MAPE is shown beside it.
  • A simple baseline won: Seasonal naive with drift beat both fitted models out of sample, so the forecast below uses it. It scored MAPE 1.27% (MASE 0.45), against seasonal naive 3.48% and naive 6.35%. The best fitted model, Damped Holt (log, deseasonalised), scored 1.65%.
  • The 80% band comes from the champion's own past errors at the same horizon: the 10th and 90th percentiles of its log errors on every checkable forecast since Jan 2012, pandemic-touching ones excluded. It is not an in-sample standard error. Checked the same way in the backtest (each band built only from errors known at the time), it held 237 of 288 times (82.3%; 95% interval 77.5% to 86.3%).
  • Seasonal decomposition: classical decomposition on logs (2x12 centred moving-average trend; seasonal factor = mean detrended value of each calendar month over its last 10 clean years; factors normalised to average zero on the log scale). The strongest month is Dec +12.0% and the weakest is Feb -10.9% against trend; the largest gap to the factors implied by Census's own adjustment (RSAFSNA / RSAFS) is 0.5 pp.
Finding
A simple baseline won, so the page uses it and says so. Its first step from Aug 2026 is -5.9% before seasonal adjustment, in line with the usual seasonal drop, and +0.55% after it, so there is no cliff. Seasonal pattern: Dec +12.0% and Feb -10.9% against trend, within 0.5 pp of Census's own factors.
Correction to the first version of this page

Correction to v1 of this page: it described FRED series RSAFS as retail sales excluding food services, not seasonally adjusted. RSAFS is retail trade and food services, seasonally adjusted, so v1's month effects were fitted to a series that had already had its seasonality removed.

v1 fitted a straight-line trend plus month effects to dollar levels over 176 months (14 years 8 months). The 4.95% error it showed was measured on the same months it was fitted to. Re-run as a forecast on the two most recent 12-month holdouts, it missed by 4.50% (Sep 2025 to Aug 2026) and 5.41% (Sep 2024 to Aug 2025), against 2.56% and 2.74% for simply repeating the last value, so it lost to the naive baseline both times.

Its first forecast month ($743.5B) sat -3.93% away from the last actual ($773.9B). The seasonally adjusted series typically moves 0.61% a month (median over the last 10 years), so that first step was a cliff, more than three typical months in one. The Forecast Lab replaces it.

Limits
  • The backtest uses today's revised history. A planner at each origin saw the advance estimates, which Census later revises, so real-world accuracy at the time would likely have been somewhat worse.
  • Trading-day and holiday effects (how many weekends a month has, when Easter falls) are not modelled; Census's seasonal adjustment accounts for them.
  • Values are nominal dollars, not adjusted for inflation.
  • The 24 scored origins cover one period (Sep 2023 to Aug 2025). The 288 scored forecasts overlap in time, so they are not independent: the band's misses fall in 8 of 35 target months, and resampling whole target months puts the coverage at about 68.3% to 93.8%, wider than the Wilson interval (77.5% to 86.3%).
  • The champion was picked on these same origins, which flatters it: re-picked at each origin from only the errors known by then, the forecasts scored MAPE 1.51%, against 1.27% for the champion chosen on all 24 origins at once.
  • The point forecast is the exponential of a log forecast, so it sits near the median rather than the mean.
  • Built from a one-off download recorded in data/fred/SOURCES.json; nothing refreshes it on a schedule. Not investment advice.
Reproduce
python tools/fetch_fred.py, then python build_forecast.py --check (script).
Seasonal factors by month: ours against those implied by Census's adjustment

Analysis D

Maintenance archetypes and a control chart

Question
Do buildings fall into a few recurring maintenance profiles across the 8 pillars? Did the citywide average evaluation score shift in any month beyond normal variation?
Method
k-means (numpy, k-means++ init, 10 restarts, seed 20260923) on the 8 pillar scores of each building's latest evaluation (0-100, unscaled; a refused area is a separate flag, a provisional choice); k from 2-6 chosen by mean silhouette on a fixed random sample of 1,500 buildings. Same k-means and silhouette search on row-centred profiles (each pillar minus the building's own mean across the 8 pillars), so clusters follow which pillar lags, not the overall level. Monthly mean PROACTIVE BUILDING SCORE of evaluations completed that month (post-2023). Centre = median of monthly means (months with n >= 20); limits = centre +/- 3 x MAD x 1.4826. Run rule: 8 or more months in a row (n >= 20) on the same side of the centre are flagged as a sustained shift.
Cluster profiles: each pillar minus the building's own average (shape clusters)
Citywide average score by month, with robust control limits
How pillar scores move together (correlation across buildings)
Finding
No distinct archetypes. On raw levels the silhouette picks 2 clusters, which only separates buildings that are good everywhere from buildings that are weak everywhere. The shape variant finds a group of 431 buildings whose paperwork pillar sits 27.6 points below their own average; 56.1% of them hold a green sign (silhouette 0.37, against 0.30 for the same data with each pillar shuffled). A one-line rule, the paperwork pillar more than 11.5 points below the building's own average across the 8 pillars, agrees with that group on 99.7% of buildings, so it is a single weak pillar, not a type of building. The control chart flags 9 of 21 eligible months (months with enough evaluations to judge). Survival analysis (time from green to first yellow) was not run for this version.
Limits
  • Cluster names are for the reader to assign; archetypes are descriptive, not causal.
  • Low silhouette values mean the clusters overlap; the score is reported so readers can judge.
  • Buildings missing any pillar (all items N/A or refused) are left out.
  • No distinct archetypes: the best split of the profiles (k = 2) has a silhouette of 0.37, against 0.30 for the same data with each pillar shuffled (no structure at all). The split is one pillar: a single rule (the paperwork pillar more than 11.5 points below the building's own average across the 8 pillars) reproduces 99.7% of it.
  • Evaluations are scheduled, not random: the mix of buildings evaluated changes by month, so a flag is a prompt to look, not evidence of a citywide change.
  • Months with fewer than 20 evaluations are shown but not flagged; 21 of the 25 months have at least 20.
  • The limits come from month-to-month swings that already include the change in which buildings are due, so they are wide (plus or minus 8.6 points) and no eligible month (20 or more evaluations) crosses them.
  • The run rule catches what the limits miss: 2023-07 to 2024-10 (9 months) all sit below the centre. The yearly averages moved from 87.2 in 2024 to 91.1 in 2025.
Reproduce
The same command as Analysis A: python build_rentsafe.py --check --input-dir <folder>, with the 4 City files from Receipts in that folder.

NorthLedger, the engine

One recorded run: the City's files through the engine

NorthLedger is the audit engine behind the Data Health Audit. It lands a file, scans it for personal data, asks a person to decide what to keep, scores its health, cleans it with every rejected row quarantined and counted, records each finding as a re-runnable fact, gates it (RECOMMEND, WATCH or INSUFFICIENT), writes the brief and exports a Power BI project. Below is one recorded run on the three City files: 22,052 rows in 3 files landed, decided and audited in 31.0 s; the post-2023 file was then analysed (422 facts, 422 refit or re-run, 422 reproduced) in 13.4 s and exported to Power BI in 10.7 s; the tidy evaluation view ran through the same stages in 11.1 s. The 17 stages sum to 66.2 s, the time the replay below plays; the whole run took 66.3 s by the wall clock, peak memory 141 MB.

Every stage of the recorded run with its measured seconds
EngagementStageRowsSecondsPeak memoryResultReplay progress
eval-2023land6,6810.55 s82 MBok
eval-2023decide00.07 s15 MBok
eval-2023audit6,68111.95 s136 MBok
eval-2023analyze6,68113.42 s141 MBok
eval-2023export6,68110.67 s130 MBok
eval-pre2023land11,7600.62 s79 MBok
eval-pre2023decide00.12 s15 MBok
eval-pre2023audit11,76011.16 s136 MBok
registrationland3,6110.44 s77 MBok
registrationdecide00.23 s15 MBok
registrationaudit3,6115.90 s135 MBok
view-2023prepare6,6810.67 s110 MBok
view-2023land6,6810.32 s68 MBok
view-2023decide00.04 s15 MBok
view-2023audit6,6813.03 s116 MBok
view-2023analyze6,6813.79 s119 MBok
view-2023export6,6813.22 s124 MBok

One recorded run on one machine (arm64, 10 CPU cores, 16 GB memory, Python 3.9.6), on 2026-09-26 12:40:32 UTC; other work was running on the same machine, so the seconds are indicative, not a guarantee. The replay plays these recorded seconds; it does not simulate typing. Engine code at the timed run: 19ec81d19e0d; engine code at this build: 19ec81d19e0d, the same code.

6 real cells, before and after

Real cells from the City's files and what happened to each
ColumnAs publishedAfter the audit
YEAR EVALUATED
evaluations 2023 to now
"202510"evaluation year 2025, taken from EVALUATION COMPLETED ON (2025-10-22); the label is flagged, not dropped
EXTERIOR GROUNDS
evaluations 2023 to now
" 2""2": whitespace trimmed, and the engine counted the repair
ELECTRICAL SAFETY PLAN
evaluations 2023 to now
"0"refused or blocked: kept as a separate flag in pillar cells (provisional); the City's own score counts it as zero points
RETAINING WALLS
evaluations 2023 to now
"N/A"not applicable by design (the building has none): left out of the score, not counted as missing
LATITUDE / X
evaluations 2023 to now
"" / "321187.241"kept: the ward code places the building; projected X/Y are not converted
SITE_ADDRESS
evaluations before 2023
"** CREATED IN ERROR ** [street address withheld here]"quarantined with its reason, never silently dropped

Found in the raw City files by rule at build time, not picked by hand.

What the run delivers

How we caught our own mistakes

  • The cleaner threw away almost everything. The first end-to-end test of a deliberately messy file kept 41 of 48,384 rows (99.9% quarantined) and still gated that finding RECOMMEND. A date rule fired on a single date-like value. Fixed with a regression test proven red first: now 46,455 clean and 1,929 quarantined (4.0%).
  • An AI analyst's total did not add up. A replayed session stated 2,564,854; the parts it had queried add to 2,565,854 (892,313 + 964 + 1,672,577). The replay page now carries the correction, and the engine's findings carry no hand-passed values.
  • The first forecast on this page lost to a naive guess. Its 4.95% error was measured on its own fitting months; out of sample it missed by 4.50% against 2.56% for repeating the last value. The Forecast Lab replaces it.
  • An earlier traceability claim was withdrawn. The replay audit now reports what it measured: 679 of 931 stated numbers (72.9%) matched a query result exactly, and 92 of 92 queries reproduced.

The messy fixture is built from Amazon Reviews'23 rows whose licence is not declared; it is used as a test input only and none of its rows are published.

Can the workflow do everything?

The honest answer, stage by stage

Mechanically yes, insightful not yet. On the messy reviews file the engine now goes from 48,384 messy rows to a cleaned table, a checked story and a backtested forecast in 60.5 s (before the engine work it kept 41 rows). On the RentSafeTO files everything it printed reproduced, but the brief is not one a landlord could use: the landlord answers (ward and pillar scorecard, points per fix, risk of losing green) still come from the site's own build script, not from the engine.

Today: messy file -> verified audit -> gated brief -> backtested forecast (generic measures only). Not yet:

  • domain measures (a weighted score, a band share, a ward x pillar table)
  • panel-aware comparisons (same building, two evaluations)
  • source-defined sentinels (a '0' that means refused) and structural N/A (no pool, no elevator)
  • a plain-language bottom line that ranks, not lists
  • charts
  • a slot that fetches and date-stamps Toronto open data
Each engine stage on three inputs, with its status and measured seconds
StageMessy file, before the fixesMessy file, afterRentSafeTO files
Land and personal-data intakeworks 7.67 sworks 7.40 sworks 1.60 s
3 files, 22,052 rows landed; 10 columns flagged across the three, 5 of them holding no personal data (item scores named 'contact' or 'tenant', a text band, a parking description; see decisions.json)
Human decisions (personal data)works works 0.69 sworks
every flagged column decided with the decide command and a written note; the rationale is in decisions.json
Source-defined sentinel decision (the RentSafeTO '0' = refused)missing missing missing
the engine has no step for it yet: the raw run averages 0 as a score (the City basis); the separate-flag decision had to be applied upstream in prepare_view.py
Profile / Data Healthpartial
Y/N booleans and pre-today typo dates missed
works 13.04 s
Health 92.7/100; profiling is still the slowest analysis sub-stage
partial 11.95 s
Health 88.2/100 on the post-2023 file, but structural 'N/A' (no pool, no elevator) is scored as missing data: 'the pools column is 93.4% null-like' is a RECOMMEND (the engine has no structural N/A step yet)
Clean + quarantinebroken
41 of 48,384 rows clean, 99.915% quarantined, gated RECOMMEND
works 2.94 s
46,455 clean + 1,929 quarantined (3.987%), gated WATCH; monthly counts a median -1.03% from the uncorrupted source over the last 24 months (before: -99.76%)
partial
post-2023: 0 quarantined, 300 cells trimmed; pre-2023: 1 quarantined (a series-end rule caught one row dated after the archive tails off); the domain defects (6 more '** CREATED IN ERROR **' rows, 2 duplicate RSN + date rows, the '202510' year label, 29 post-2023 year labels that differ from the completion date) are caught only by build_rentsafe.py
Latest row per building (panel collapse)missing partial partial
clean_table(latest_per_key=RSN) keeps 3,588 rows, the same rows as the site build (identical); the CLI has no option to ask for it
Facts, gate, verifyworks
21 of 21 SQL facts reproduced; the 20 cleaning facts were not re-runnable
works 0.30 s
80 of 80 analysis facts reproduced on re-run or refit
works
422 of 422 analysis facts reproduced (raw file); 158 of 158 (view)
Business measuresmissing works 0.56 s
volume -32.4% (uncorrupted source -32.8%), star rating +0.141 (source +0.140)
partial
on the raw file 56 columns were measured and 12 left out, each with its stated reason: id (an identifier, not a measure); rsn, ward (a code or key, not a quantity); year_registered, year_built, year_evaluated (a year label, not a quantity); site_address (personal data decision at intake); grid (free text or too many distinct values to be a category); latitude, longitude, x, y (a map coordinate, not a measure); the engine's 12-month comparison is not panel-aware: its windows compare mostly different buildings (268 of 2,043 in both), so the engine's +0.77-point row-weighted score change (its headline: +0.7% in the average month) hides a same-building change of +3.64 points
Forecast with backtestmissing
prototype only: naive MAPE 19.1%, 80% band held 0.67
works 0.06 s
champion holt_damped_log, INSUFFICIENT; replayed 24 months, month-ahead band held 18 of 24
works (honestly declines)
monthly evaluations forecast gated INSUFFICIENT: only 6 replay months in 39 months of history, and 98.9% of evaluations fall June to November, so most months are empty by design; on RSAFSNA the engine's champion is seasonal_naive (baseline won) while the Forecast Lab's is snaive_drift
Story: what happened / why / what's next / what to domissing
the brief told the client to fix 99.92% quarantined rows
partial 0.00 s
numbers are cited facts, but 'Why' is empty (the end of collection after 2023-08 is not named as a cause) and 'What to do' recommends acting on helpful_votes
partial
the raw bottom line lists 45 changes in one sentence, unranked (it opens with average abandoned_equip_derelict_veh, average building_cleanliness, average building_exterior), and its 'Do now' recommends no action: the change test could not run on any of the 58 measured changes (57 of them because too few months in a 12-month window hold a value: evaluations cluster in part of the year); on the view a +2.9% change in paperwork_score, averaged by month, had no change test (only 7 of the latest 12 months and 5 of the 12 before hold a usable value; the test needs at least 10 in each; the file holds no values for January, February, March, April, May in any year), and the brief states that reason; no charts; no ward or pillar comparison
Power BI export (PBIP)partial 15.36 s
41 rows, empty page, SUM-of-ratings measures
works 16.69 s
lint clean; Forecast and Backtest tables
works (not opened in Desktop) 10.67 s
pbip_lint clean; Microsoft TMDL parser loaded 6 tables and 1 relationship (raw) and 6 tables (view); never opened in Power BI Desktop
Toronto open-data enrichment slotmissing missing missing
files were downloaded once by hand (SOURCES.json); the engine cannot fetch or date-stamp them

The same-building change in the Business measures row (+3.6 points over 3,093 pairs), looked at more closely. Split by each pair's first score: pairs that started under 70 moved +22.5 points on average (126 pairs, 96.0% higher), and pairs that started at 95 and over moved -1.8 (596 pairs, 32.2% higher). Low starters rising while high starters fall is the pattern regression toward the middle produces (the 100-point cap also limits high starters), so part of the average may not be repair.

The messy file, after the fixes

48,384 messy rows: land to Power BI export, analysis included, in 60.5 s. 48,384 rows in = 46,455 clean + 1,929 quarantined. Against the uncorrupted source, the engine's volume change is -32.4% (source -32.8%) and its rating change +0.141 (source +0.140). Its forecast missed by 19.8% one month ahead (52.6% over all horizons), and its month-ahead band held 18 of 24 replayed months.

Two forecasters, named apart

On the same retail file, the Forecast Lab picks snaive_drift and the engine picks seasonal_naive. Where they share a method they agree exactly. The differences, named:

  • Model set: the lab's champion, seasonal naive with drift ('snaive_drift'), does not exist in the engine, whose five models are naive, seasonal naive, drift, damped Holt on logs and log-linear + season.
  • Holt: the lab fits damped Holt on the deseasonalised log series; the engine's Holt has no seasonal term, so on a strongly seasonal, not-seasonally-adjusted series it cannot track December and February.
  • Pandemic window: the lab treats Mar 2020 - Jun 2021 as missing when it estimates drift, seasonality, Holt and the log-linear model; the engine fits through it (its log-linear window is the last 60 months, which at the lab's first origins still contains 2020-21).
  • Backtest design: the lab scores 24 origins (Sep 2023 - Aug 2025) x all 12 horizons = 288 forecasts per model; the engine's 24 origins end at the last month, so later origins score fewer horizons (a triangle weighted towards short horizons), and it uses only the last 120 months.
  • Error scale: MASE uses a different denominator in each code (the lab: 120-month seasonal-naive MAE before each origin, pandemic pairs excluded; the engine: seasonal-naive MAE before the first scored origin), so only MAPE is compared across codes.
  • Every timing is one recorded run on one machine (see machine); other work was running on the same machine.
  • The status words are judgments; the evidence beside each is computed from the named files.
  • The PBIP files have never been opened in Power BI Desktop; only pbip_lint and Microsoft's TMDL parser checked them.
  • The messy fixture is built from Amazon Reviews'23 rows whose licence is not declared; it is used as a test input only and none of its rows are published.

Try it on your own file

Your messy CSV, through the same engine, inside your browser

Drop a spreadsheet export and the NorthLedger engine from the run above works on it here, on your computer. It cleans the file with stated rules and counts every fix, sets aside the rows it cannot trust and says why, grades each finding (CONFIRMED, WATCH or NOT ENOUGH DATA), backtests a forecast when the file has a date column, and writes the story. You can download the cleaned file and the evidence afterwards.

The engine holds back on small files, on purpose: a forecast needs at least 45 months of monthly history, so that the last 12 can be replayed to test it, and a finding must rest on 1,000 rows before the engine will grade it CONFIRMED. A smaller file still gets the cleaning, the data-health check and the checked facts.

Drop a CSV or TSV file here

or

Up to 25 MB and 200,000 rows. A larger file is refused, never sampled.

No file to hand?

The sample is a deliberately messy rent and expense ledger for five invented Toronto buildings, made for this demo: no real people or accounts in it.

Type it before you drop the file. If you then choose "Continue with the AI", the AI reads a summary of your columns (never rows, and never a column you withhold), plans the analyses around your question, and the engine computes every number in your browser. Leave it empty and the AI proposes the goal itself; the plan also offers other questions this data can answer, one click each.

Runs in your browser. Your file is not uploaded anywhere. The first run fetches the engine's Python runtime (Pyodide) from this site, and your file never goes with it. The AI is used only if you choose "Continue with the AI". It plans from a summary of your columns, never rows; a column you withhold is never sent or named, and a personal column you keep is sent only if you tick the box that names it. The engine runs the plan and computes every figure; the AI corrects its plan once if the engine finds a problem, then writes the report from the engine's results.

The AI analyst, replayed

Recorded sessions, played back as they happened

11 recorded sessions of an AI data analyst working on public data: 10 featured, 1 archived. These recorded sessions send query results to a remote model (deepseek-chat, glm-5.3). The Data Health Audit offered below is a different tool: the NorthLedger engine runs it, and its runs recorded on this page made no AI model calls.

  • AI-agent data-quality session (messy legacy export)

    client_data.db (demo database)

    17 turns, 29 tool calls; report: PDF

    Open the replays
  • Forecast + risk brief

    client_data.db (demo database)

    25 turns, 37 tool calls; report: DOCX

    Open the replays
  • Dutch fleet electrification (16.9M vehicles)

    big_data.db (scale tier)

    9 turns, 28 tool calls; report: PDF

    Open the replays

Replay audit: every recorded query was re-run on the same data, and 92 of 92 reproduced (same query, same data; that proves reproducibility, not that each query asked the right question). 679 of 931 numbers stated in the sessions matched a query result exactly (72.9%). Editor's notes on the replay page are marked as added after the session.

Power BI, AI-built and owner-directed

Power BI project: NYC 311 Service Requests (PL-300 skill areas)

dev profile refreshed and reconciled in the Power BI service; full profile, RLS and report pages not yet run

AI-built, owner-directed: the Power Query, model, measures, RLS and report were written by AI (Claude) from a specification Rashadul directed.

The numbers are checked against an independent Python twin with the reconcile CLI. On the dev profile every default check matched the twin cell for cell. The full profile has not been refreshed, so no figure from it is claimed.

It runs in the Power BI service in the browser: Fabric Git integration syncs the project from GitHub into a Premium Per User trial workspace, where the dev profile, a sample of the data, has been refreshed.

Next, on the owner's side: refresh the full profile and reconcile it, test the RLS roles with Test as role, and review the report pages and add real screenshots.

The repository holds the model, report, check queries, reconcile CLI and specification, and the dev profile's check results and reconcile record; the release holds the full data extract the model reads, published as a release because it is too large for the repository. Both are public.

Milestones

  • Data model and Power Query refreshed, dev profile
  • Measures (DAX) checked on the dev profile
  • Project files and full data extract, public on GitHub published
  • Power BI workspace (Premium Per User trial), synced by Fabric Git set up
  • Refresh and reconcile against the Python twin, dev profile passed
  • Refresh and reconcile, full profile not yet run
  • Row-level security, tested with Test as role not yet run
  • Report pages reviewed, with real screenshots not yet
  • Publish to web (screenshot or embed) not yet

The reconcile record is kept in the project itself: the dev refresh and reconciliation record. This page does not copy its figures. This card reads only pl300-status.json and claims nothing that has not been run: no figure from this model appears here until the owner records one.

Data: NYC 311 Service Requests from 2010 to Present (NYC Open Data).

Take it into Power BI

Your audited data arrives as a Power BI project, not only a PDF

The engine's export writes a Power BI Project (PBIP): the semantic model in TMDL text files, a report shell and the cleaned data as CSV. Here it is for the RentSafeTO evaluation view: 6 tables, 26 engine-generated measures, 1 relationship, 26 files.

What it holds, and what it does not: every current-method evaluation, one row each (6,681 rows, 2023-06-05 to 2026-09-21, every evaluation cycle), not the report's latest evaluation per building (3,588 buildings). So its averages differ from the cards above: its mean current score is 88.9, against a simple mean of 90.4 in the report. It has no building key and no unit count, so the report's latest-per-building view and its unit-weighted score cannot be rebuilt from it yet.

Status, at exactly the level reached: structure checked on a Mac with pbip_lint (clean); the TMDL loaded by Microsoft's own parser (powerbi-modeling-mcp 0.5.0.0: 6 tables, 1 relationship); DAX and Power Query not evaluated; not yet opened in Power BI Desktop. Power BI Desktop runs only on Windows, and PBIP is a Microsoft preview feature.

Inside the project

Click a highlighted file to read its first lines.

  • .gitignore 45 B
  • README.md 3 KB
  • 292 B
  • RentSafeTO Scorecard View.Report/
  • .platform 310 B
  • definition.pbir 252 B
  • definition/
  • report.json 144 B
  • version.json 150 B
  • pages/
  • pages.json 219 B
  • bbc221e497a29dbc9b15/
  • page.json 248 B
  • RentSafeTO Scorecard View.SemanticModel/
  • .platform 317 B
  • definition.pbism 167 B
  • definition/
  • database.tmdl 35 B
  • expressions.tmdl 275 B
  • model.tmdl 313 B
  • 134 B
  • tables/
  • Backtest.tmdl 2 KB
  • Data Quality.tmdl 2 KB
  • 2 KB
  • 3 KB
  • Quarantine.tmdl 3 KB
  • 6 KB
  • data/
  • backtest.csv 1 KB
  • clean.csv 964 KB
  • data_quality.csv 39 KB
  • forecast.csv 2 KB
  • quarantine.csv 264 B

Download the PBIP (zip, 161 KB) checksum starts d3a4c97c66caccbf

Generated by the NorthLedger engine (AI-built). Open it on Windows, set the DataFolder parameter to the unzipped data folder, refresh, and check the row counts in its README.

For business owners

Your data, handled, with receipts

Most small businesses do not need a data scientist on payroll. They need their exports and spreadsheets to be cleaned, to agree with each other, and to say something useful. Every finding comes with the query that produced it, and anything I could not check is labelled as such.

  1. CollectYour exports and spreadsheets, landed with a record of what arrived.Built, public data only
  2. CleanDuplicates, formats and gaps fixed by tested rules; rejects quarantined, never silently dropped.Built + tested
  3. ModelMeasures defined against your objectives, exported as a Power BI project.Built, partial
  4. InsightA written brief that only states numbers a check has re-run.Built, partial
  5. ReportAn interactive report like the one above, rebuilt with one command when new data arrives.Built, public data only
  6. PredictBacktested forecasts with honest bands; the simple baseline wins when it should.Built + tested

Which offer fits?

Three questions; nothing you pick leaves this page.

How many sources?
How often?
Forecasts?

Entry

Data Health Audit

A fixed-scope diagnosis of what you have: what is broken, what can be automated, and the three changes that pay for themselves first. You get a written verdict whether or not we continue.

From $500 CAD, fixed 1 to 2 weeks. Delivered as a scored report and a 30-minute call.

AI model calls in the audit runs recorded on this page: none.

The build

Automated Insights Build

The full loop on your data: sources connected, cleaned and modelled against your objectives, then a report, a written brief and a backtested forecast. Fixed scope, handover runbook included.

From $4,000 CAD, setup 4 to 8 weeks, scaled to source count. Two training sessions included.

Operate

Insights Retainer

The system keeps working: monitoring, monthly summaries, forecast refreshes, one deep-dive analysis a month (like the gallery cards on this page), and small enhancements as the business changes.

From $800 CAD / month Ongoing. Your data person on call, without the payroll line.

What you get, shown on real samples

Start with an audit

Book

Start with an audit

One call about your messiest spreadsheet, and you leave with a verdict: what is broken, what can be automated, and what it pays back.

Research and academic collaboration: roman4@uwindsor.ca.

Please don't attach data to a first message. After a short call I send a private upload link.

Running a research, open-source or data-team project instead? Research collaboration

  1. CallWhat you want to know, and which tools hold the data.
  2. Private uploadA link just for you; nothing goes through email.
  3. VerdictA scored report, a written brief and a short walkthrough call.
  4. You keep itThe report is yours, with or without a retainer.

Who is behind this

I'm Rashadul Islam Roman, a data analyst and AI-automation instructor based in Toronto. My method is evidence-first: no claim ships without a source, and no number without a way to regenerate it. I teach professionals to build automations for a living, and I hold my own work to the standard I teach.

How this page was made: the engine, the analyses and this site were written with AI assistance (Hermes Agent with GLM, and Claude) and checked by the tests and build checks linked below. The PL-300 Power BI project was also written by AI, from a specification the owner directed. Labels on each artifact say which is which.

  • MSc Applied Economics, University of Windsor
  • Based in Toronto, working with small and mid-size businesses
  • Automation practice: 76+ production automations built and taught across five business domains (the owner's own statement; not checked here)
  • Tools used on this page: Python (pandas, numpy), SQL (SQLite), Power BI project files (TMDL), hand-written SVG. Not used here yet: BigQuery, dbt, Airflow, gradient-boosted models Roadmap, not built

Research collaboration

Invite me to a project

I contribute to research and data projects as an independent analyst. If your project needs careful data work, with a written record of how every number was made, send a short invitation and I will reply with whether and how I can help.

Who it is for

  • Researchers and labsData cleaning, analysis and reproducible pipelines for a study, a report or a working paper.
  • Open-source and data teamsData engineering, tests and documentation on shared code and public datasets.
  • BusinessesA scoped analysis or forecast, worked alongside your own team, with its method written down.

What I contribute

  • Data engineering and cleaningMessy exports turned into tested, documented tables; rows that fail a rule are set aside with the reason, never dropped silently.
  • Statistical analysis and forecastingTests reported with effect sizes and intervals; forecasts backtested out of sample, with the measured error rate stated beside a simple baseline.
  • Power BI and dashboardsSemantic models, measures and reports, reconciled against an independent calculation before a number is claimed.
  • AutomationScripts and pipelines that rebuild a result from its inputs with one command.

How the engine's trust rules apply to shared work

  • Every number is computed by code kept with the project, so anyone on the team can regenerate it.
  • Each claim carries a verdict: a result not yet tested against chance is labelled as such, not reported as a finding.
  • Rows that fail a check are quarantined with the reason and counted, never deleted quietly.
  • Personal data stays withheld unless the question needs it; your data agreement and any ethics approval set the terms.
  • AI-assisted work is labelled as AI-assisted, as it is on this page.
Invite me to a project

The button opens an email to roman4@uwindsor.ca with the subject line and a few short prompts filled in: the project, the role, the timeline, the data involved and links. Please don't attach data to a first message. For business enquiries, write to contact@halosyncs.com.

No institutional affiliation is claimed or implied: I work as an independent analyst in Toronto. What I can show is on this page, with its receipts.

Receipts

Sources, build and errata

Every tile and claim above links here. Nothing on this site runs on a schedule or fetches data when you visit: each dataset is a download someone made on purpose, recorded with its address, time, size and checksum.

Data sources and licences

Every downloaded input with its retrieval time, size and checksum
Dataset and fileWhere fromRetrievedSizeChecksum, first charactersLicence
Apartment Building Evaluation
apartment-building-evaluations-2023-current.csv
CKAN download2026-09-23 06:40:59 UTC1.8 MBdb9feec11327ca41Open Government Licence, Toronto
Apartment Building Evaluation
pre-2023-apartment-building-evaluations.csv
CKAN download2026-09-23 06:40:59 UTC2.8 MB9ddfb28db7d35453Open Government Licence, Toronto
Apartment Building Registration
apartment-building-registration-data.csv
CKAN download2026-09-23 06:40:59 UTC1.6 MB9ca50a07cd567d68Open Government Licence, Toronto
City Wards
city-wards-data-4326.geojson
CKAN download2026-09-23 06:41:00 UTC1.1 MBa35851f39c83e492Open Government Licence, Toronto, assumed: the catalogue entry names none
U.S. retail trade and food services, RSAFSNA
U.S. Census Bureau, via FRED (Federal Reserve Bank of St. Louis)
FRED series page2026-09-23 06:47:06 UTC
source file dated Wed, 16 Sep 2026 12:39:52 GMT
8 KB2bbe15d2fc830b6bFRED terms, with attribution
U.S. retail trade and food services, seasonally adjusted, RSAFS
U.S. Census Bureau, via FRED (Federal Reserve Bank of St. Louis)
FRED series page2026-09-23 06:47:06 UTC
source file dated Wed, 16 Sep 2026 12:39:53 GMT
8 KBc57685eed78fe9d2FRED terms, with attribution

City of Toronto data. Contains information licensed under the Open Government Licence – Toronto. (the licence). This is independent analysis by NorthLedger Insights; it is not produced by, affiliated with or endorsed by the City of Toronto or RentSafeTO. The sign bands (green, yellow, red) are the City's own, from its page on the colour-coded signs; whether signs issued for older evaluations used the same cut-offs was not checked for this build. The ward boundaries come from the City Wards package, whose catalogue entry names no licence; the portal-wide licence is assumed there.

Retail sales. U.S. Census Bureau, via FRED (Federal Reserve Bank of St. Louis). Series RSAFSNA and RSAFS. FRED's terms for redistributing derived charts were not re-read for this build.

Design credit. The scorecard's layout is inspired by the WEF Travel and Tourism Development Index table; no TTDI data is reproduced.

Build and tests

Errata, published rather than fixed silently

  • The first version of this page mislabelled its retail series and showed an in-sample error as accuracy; its forecast lost to a naive guess out of sample. See the correction in Analysis C.
  • A replayed AI session stated 2,564,854 where its own parts add to 2,565,854.
  • The engine's cleaner once kept 41 of 48,384 rows of a messy file; fixed, with the test that caught it.
  • An earlier traceability figure for the replays was withdrawn; the measured figure is 72.9%.

Accessibility and privacy

Target: WCAG 2.2 AA. Every chart has a text name and a table view, every focusable chart mark has a spoken name, tooltips open on keyboard focus and close with Esc, the scorecard is a real table on every screen size, and text contrast is checked at build (worst pair 4.76:1). Behaviour is checked in a headless browser at build time (tools/check_ui.js). Known gaps: the charts need JavaScript (their figures are in the text and in the data files without it), and the page has not been tested with a screen reader by a person.

No trackers and no analytics. While you read, the page makes no request to another address. Starting the Try-it demo fetches the engine's Python runtime (Pyodide) and the engine itself from this site; only if you choose "Continue with the AI", it sends a summary of your columns (never rows, never a column you withhold) and then the engine's results to the owner's AI proxy, which passes them to DeepSeek. Your file itself is never sent anywhere.