{"question":"Will a building's next biennial evaluation score fall below 85 (lose or miss the green sign)?","target":"next PROACTIVE BUILDING SCORE < 85. The proactive score is used because reactive deductions are only recorded on each building's latest row (a refresh-time snapshot), so they are not available at evaluation time for earlier rows.","pairs":{"total":3093,"train":1721,"test":1372,"split":"temporal: pairs whose NEXT evaluation completed before 2026-01-01 train; on/after test","train_positive_rate":0.158,"test_positive_rate":0.1582,"bottom_2_5pct_pairs_all":88},"features":["last_score","pillar_envelope","pillar_grounds","pillar_entry","pillar_systems","pillar_stairs","pillar_interiors","pillar_waste","pillar_paperwork","refused_items","year_built","log_storeys","log_units","tchc","social","elevator_original","heating_original","single_pane","log_firm_portfolio","district_north","district_toronto","district_scarborough","pre2023_last_pctile","pctile_change_vs_pre2023","has_pre2023"],"feature_notes":"From the earlier evaluation of each pair plus registration attributes. Pillar scores keep a refused area as a separate flag (a provisional choice, to be confirmed). Registration is a single 2026-07-05 snapshot, so elevator/heating/window/firm fields may postdate the earlier evaluation (a leakage risk stated as a limit). Firm portfolio size = buildings per normalised management-firm name (names used only inside the model; never published).","model":{"logistic":"numpy Newton-Raphson logistic regression, L2 lambda=100 (chosen on an inner temporal split of the training pairs at 2025-08-01)","lambda_search":[{"lambda":0.1,"val_brier":0.14018},{"lambda":1,"val_brier":0.13796},{"lambda":3,"val_brier":0.13539},{"lambda":10,"val_brier":0.13096},{"lambda":30,"val_brier":0.1269},{"lambda":100,"val_brier":0.12472},{"lambda":300,"val_brier":0.12608}],"score_model":"numpy ridge regression, lambda=300","ridge_lambda_search":[{"lambda":0.1,"val_mae":4.1884},{"lambda":1,"val_mae":4.1889},{"lambda":3,"val_mae":4.1893},{"lambda":10,"val_mae":4.1889},{"lambda":30,"val_mae":4.188},{"lambda":100,"val_mae":4.1833},{"lambda":300,"val_mae":4.1823}],"coefficients_per_sd":[{"feature":"pre2023_last_pctile","coef_per_sd":-0.3289},{"feature":"pillar_interiors","coef_per_sd":-0.2477},{"feature":"pillar_systems","coef_per_sd":-0.2156},{"feature":"log_firm_portfolio","coef_per_sd":-0.1335},{"feature":"last_score","coef_per_sd":-0.113},{"feature":"tchc","coef_per_sd":0.1096},{"feature":"year_built","coef_per_sd":-0.1051},{"feature":"pillar_grounds","coef_per_sd":-0.1011},{"feature":"district_toronto","coef_per_sd":0.0838},{"feature":"pillar_entry","coef_per_sd":-0.0693},{"feature":"heating_original","coef_per_sd":-0.0634},{"feature":"pillar_waste","coef_per_sd":-0.0608},{"feature":"pillar_stairs","coef_per_sd":-0.0564},{"feature":"log_storeys","coef_per_sd":0.0449},{"feature":"has_pre2023","coef_per_sd":0.0422},{"feature":"social","coef_per_sd":0.0322},{"feature":"log_units","coef_per_sd":-0.0259},{"feature":"district_north","coef_per_sd":0.0257},{"feature":"elevator_original","coef_per_sd":0.0197},{"feature":"single_pane","coef_per_sd":-0.0142},{"feature":"pillar_paperwork","coef_per_sd":-0.0076},{"feature":"pillar_envelope","coef_per_sd":0.0073},{"feature":"refused_items","coef_per_sd":0.0045},{"feature":"pctile_change_vs_pre2023","coef_per_sd":0.0026},{"feature":"district_scarborough","coef_per_sd":0.0019}]},"baselines":{"persistence_band_rate":"P(next < 85) = training-set rate of next < 85 among pairs whose last score is in the same band [0, 70, 80, 85, 88, 91, 94, 101]","persistence_band_rates":{"0-69":0.3302,"70-79":0.4167,"80-84":0.2449,"85-87":0.1908,"88-90":0.1097,"91-93":0.0892,"94-100":0.0359},"naive_last_value":"predict next score = last score; P(next < 85) = 1 if last < 85 else 0"},"test_metrics":{"brier_model":0.12487,"brier_persistence_band":0.12115,"brier_naive_0_1":0.28207,"pr_auc_model":0.3292,"pr_auc_persistence_band":0.308,"pr_auc_no_skill":0.1582,"mae_score_model":4.9665,"mae_score_persistence":7.1079,"mae_score_persistence_plus_band_mean_change":5.0848,"mae_note":"Plain persistence (next = last) ignores the typical rise between evaluations; the band-mean-change baseline adds the training mean change for the last-score band. The gate uses plain persistence as specified; the stronger baseline is shown for honesty.","interval80_model":{"lo_offset":-7.56,"hi_offset":5.552,"coverage_test":0.7172,"k":984,"n":1372,"wilson95":[0.6928,0.7404],"method":"offsets = 10th/90th percentiles of training residuals"},"interval80_persistence":{"lo_offset":-5.0,"hi_offset":13.0,"coverage_test":0.7784}},"reliability_model":[{"bin":"0.0-0.1","n":584,"mean_predicted":0.0657,"observed_rate":0.0548},{"bin":"0.1-0.2","n":434,"mean_predicted":0.1407,"observed_rate":0.1636},{"bin":"0.2-0.3","n":160,"mean_predicted":0.2424,"observed_rate":0.25},{"bin":"0.3-0.4","n":91,"mean_predicted":0.3453,"observed_rate":0.3956},{"bin":"0.4-0.5","n":38,"mean_predicted":0.4501,"observed_rate":0.3158},{"bin":"0.5-0.6","n":23,"mean_predicted":0.5418,"observed_rate":0.4348},{"bin":"0.6-0.7","n":21,"mean_predicted":0.6566,"observed_rate":0.2857},{"bin":"0.7-0.8","n":7,"mean_predicted":0.7385,"observed_rate":0.4286},{"bin":"0.8-0.9","n":10,"mean_predicted":0.8404,"observed_rate":0.4},{"bin":"0.9-1.0","n":4,"mean_predicted":0.9331,"observed_rate":0.75}],"reliability_persistence":[{"bin":"0.0-0.1","n":525,"mean_predicted":0.0601,"observed_rate":0.0514},{"bin":"0.1-0.2","n":393,"mean_predicted":0.145,"observed_rate":0.1221},{"bin":"0.2-0.3","n":209,"mean_predicted":0.2449,"observed_rate":0.2632},{"bin":"0.3-0.4","n":74,"mean_predicted":0.3302,"observed_rate":0.4324},{"bin":"0.4-0.5","n":171,"mean_predicted":0.4167,"observed_rate":0.3216}],"gate":{"rule":"The model is champion only if it beats persistence on BOTH Brier (P(next<85)) and MAE (next score) out of sample.","champion":"persistence","verdict_text":"The model was not offered: it beat repeating the last score on average miss (4.97 against 7.11 points), but not the band's past rate on Brier score (0.125 against 0.121), and the pre-set rule needed both."},"by_ward_test":[{"ward":"01","n_pairs_test":21,"mean_model_risk":0.1879,"mean_persistence_risk":0.1755,"observed_rate":0.0476},{"ward":"02","n_pairs_test":57,"mean_model_risk":0.1423,"mean_persistence_risk":0.1818,"observed_rate":0.0877},{"ward":"03","n_pairs_test":90,"mean_model_risk":0.1139,"mean_persistence_risk":0.1252,"observed_rate":0.1333},{"ward":"04","n_pairs_test":83,"mean_model_risk":0.1795,"mean_persistence_risk":0.1624,"observed_rate":0.1928},{"ward":"05","n_pairs_test":110,"mean_model_risk":0.1593,"mean_persistence_risk":0.1795,"observed_rate":0.3273},{"ward":"06","n_pairs_test":97,"mean_model_risk":0.2017,"mean_persistence_risk":0.2023,"observed_rate":0.0825},{"ward":"07","n_pairs_test":30,"mean_model_risk":0.1465,"mean_persistence_risk":0.1587,"observed_rate":0.0333},{"ward":"08","n_pairs_test":99,"mean_model_risk":0.1851,"mean_persistence_risk":0.1888,"observed_rate":0.0808},{"ward":"09","n_pairs_test":34,"mean_model_risk":0.4264,"mean_persistence_risk":0.2969,"observed_rate":0.3529},{"ward":"10","n_pairs_test":27,"mean_model_risk":0.2005,"mean_persistence_risk":0.1975,"observed_rate":0.037},{"ward":"11","n_pairs_test":63,"mean_model_risk":0.2112,"mean_persistence_risk":0.1917,"observed_rate":0.1429},{"ward":"12","n_pairs_test":148,"mean_model_risk":0.1334,"mean_persistence_risk":0.1282,"observed_rate":0.0541},{"ward":"13","n_pairs_test":69,"mean_model_risk":0.2095,"mean_persistence_risk":0.2019,"observed_rate":0.2464},{"ward":"14","n_pairs_test":61,"mean_model_risk":0.1279,"mean_persistence_risk":0.1526,"observed_rate":0.3443},{"ward":"15","n_pairs_test":57,"mean_model_risk":0.1026,"mean_persistence_risk":0.1118,"observed_rate":0.0175},{"ward":"16","n_pairs_test":57,"mean_model_risk":0.1231,"mean_persistence_risk":0.1269,"observed_rate":0.0877},{"ward":"17","n_pairs_test":16,"mean_model_risk":0.1636,"mean_persistence_risk":0.2086,"observed_rate":0.3125},{"ward":"18","n_pairs_test":9,"mean_model_risk":0.0796,"mean_persistence_risk":0.1036,"observed_rate":0.2222},{"ward":"19","n_pairs_test":97,"mean_model_risk":0.2166,"mean_persistence_risk":0.2358,"observed_rate":0.1134},{"ward":"20","n_pairs_test":62,"mean_model_risk":0.1529,"mean_persistence_risk":0.1713,"observed_rate":0.2581},{"ward":"21","n_pairs_test":37,"mean_model_risk":0.1072,"mean_persistence_risk":0.1217,"observed_rate":0.2703},{"ward":"22","n_pairs_test":19,"mean_model_risk":0.1558,"mean_persistence_risk":0.1686,"observed_rate":0.0526},{"ward":"23","n_pairs_test":2,"mean_model_risk":0.0295,"mean_persistence_risk":0.0359,"observed_rate":0.0},{"ward":"24","n_pairs_test":20,"mean_model_risk":0.1503,"mean_persistence_risk":0.1516,"observed_rate":0.3},{"ward":"25","n_pairs_test":7,"mean_model_risk":0.1992,"mean_persistence_risk":0.2693,"observed_rate":0.7143}],"transition_counts_all_pairs":{"rows_last_band":["green","yellow","red"],"cols_next_band":["green","yellow","red"],"counts":[[1952,177,7],[575,246,10],[77,37,12]]},"limits":["About one post-2023 pair per building; the test set is one season (2026) and small.","Biennial evaluations: a building's condition between evaluations is not observed.","Registration attributes are a 2026 snapshot (possible look-ahead).","Reliability bins with few buildings are noisy; read them with their n."],"stage":"stage 2: NorthLedger-cleaned frames (engine 19ec81d19e0d), then this script's domain rules","data_as_of":"2026-09-21","built_at":"2026-09-26T12:43:42Z"}