Engineering Council Test Reliability Report

Scope aligned with Slack channel #dezvoltare, covering 2026-07-26 10:00 to 2026-08-02 10:00. Metrics and timings are sourced from GitLab pipelines, jobs, and test-report artifacts for the daily 6 PM regression suite and the production smoke suite. Trend charts use daily buckets across this window.

Executive Snapshot

6
Daily Runs
2/6
Daily Green
12m 46s
Avg Daily Runtime
17
Smoke Attempts
17/17
Smoke Green
4m 03s
Avg Smoke Runtime
4m 11s
Median Smoke Time
0
Current Green Streak

Executive Analysis

Bottom line: the regression system is informative but not calm. The data suggest repeatable problem areas rather than random breakage, which means focused ownership should move the needle quickly.

What Matters

  • Daily regression passed 2 of 6 runs (33.3%), with a current green streak of 0 and a best streak of 1 in this window. The latest daily run (165195) failed, so the system is ending the week under tension rather than in a clean state.
  • Smoke passed 17 of 17 attempts (100.0%) across 12 production pipelines.
  • Failure concentration is not random: Frontend has the highest strict failure ratio at 0.32%, while Billing has the broadest non-pass footprint at 30.79%.
  • Frontend is the weakest smoke surface in this window at 12/12 green (100.0%).
  • Daily-suite runtime averaged 12m 46s.

Engineering Analysis

  • The failure profile is concentrated enough to act on. Frontend and Billing are carrying the strongest signal, which means reliability work should be assigned by category ownership instead of treating the suite as one undifferentiated problem.
  • The broader daily suite is carrying more instability than smoke, which usually means product regressions are escaping into wider coverage areas even when the narrow deploy gate looks acceptable.

Recommended Actions

  • Assign one owner to Frontend for the next cycle and expect a short written burn-down: top failing tests, suspected root causes, flake versus regression breakdown, and what gets fixed or quarantined first.
  • Treat the daily regression suite like an operations queue until it is calm again: triage failures after each red run, close known-noise items fast, and avoid letting multiple unrelated red signals pile up between runs.
  • Put Frontend smoke under closer guardrails for the next release cycle. It is the best place to improve first-pass deploy confidence quickly.

Improvement Ideas

  • Introduce a small reliability budget for tests: every flaky or quarantined case needs an owner and an expiry, and the team should review that budget weekly the same way it reviews bugs or incidents.
  • Track first-fail to root-cause time as a core metric. Fast diagnosis is as important as raw pass rate because the practical value of a test gate depends on how quickly it helps the team recover.
  • Define a runtime budget per suite and require justification when test count or duration grows. Reliable feedback systems stay trusted when they remain both stable and proportionate.

Category Execution Ratios

How computed

Category total executions means the sum of that category's observed test executions across every daily-suite run in the selected window.

Strict Failure Ratio = failed executions for that category divided by total executions for that category across the window.

Non-pass Ratio = (failed + pending + skipped) executions for that category divided by total executions for that category across the window.

Example: if Billing executed 800 times across the week and 2 of those executions failed, Billing strict failure ratio is 0.25%. That does not mean 0.25% of pipelines failed; it means 0.25% of observed Billing executions ended in failed.

How computed

Category total executions means the sum of that category's observed test executions across every daily-suite run in the selected window.

Strict Failure Ratio = failed executions for that category divided by total executions for that category across the window.

Non-pass Ratio = (failed + pending + skipped) executions for that category divided by total executions for that category across the window.

Example: if Billing executed 800 times across the week and 2 of those executions failed, Billing strict failure ratio is 0.25%. That does not mean 0.25% of pipelines failed; it means 0.25% of observed Billing executions ended in failed.

Daily Daily Suite Status0000107-2607-2707-2807-2907-3007-31
Daily Smoke Attempts0123507-2607-2707-2807-2907-3007-31
Daily Average Daily Suite Runtime12m 06s12m 24s12m 41s12m 59s13m 17s07-2607-2707-2807-2907-3007-31
Daily Average Smoke Runtime0m 00s1m 08s2m 16s3m 24s4m 32s07-2607-2707-2807-2907-3007-31
Daily Suite Total Test Growth (Recent 6 Runs)1503150315031503150407-2607-2707-2807-2907-3007-31
Smoke Suite Total Test Growth (Latest Run Per Day)
FrontendUniversity
60728597110Frontend 07-27: 110Frontend 07-28: 110Frontend 07-30: 110Frontend 07-31: 110University 07-29: 60University 07-30: 60University 07-31: 6007-2707-2807-2907-3007-31

Category Aggregate Table

How computed

Category total executions means the sum of that category's observed test executions across every daily-suite run in the selected window.

Strict Failure Ratio = failed executions for that category divided by total executions for that category across the window.

Non-pass Ratio = (failed + pending + skipped) executions for that category divided by total executions for that category across the window.

Example: if Billing executed 800 times across the week and 2 of those executions failed, Billing strict failure ratio is 0.25%. That does not mean 0.25% of pipelines failed; it means 0.25% of observed Billing executions ended in failed.

How computed

Category total executions means the sum of that category's observed test executions across every daily-suite run in the selected window.

Strict Failure Ratio = failed executions for that category divided by total executions for that category across the window.

Non-pass Ratio = (failed + pending + skipped) executions for that category divided by total executions for that category across the window.

Example: if Billing executed 800 times across the week and 2 of those executions failed, Billing strict failure ratio is 0.25%. That does not mean 0.25% of pipelines failed; it means 0.25% of observed Billing executions ended in failed.

CategoryTotalFailedPendingSkippedFailure RatioNon-pass RatioRuns With Failures
Billing786224000.25%30.79%2
Web46682020.04%0.09%2
Frontend19026000.32%0.32%2
Library5160000.00%0.00%0
University180000.00%0.00%0
Subscriptions60000.00%0.00%0
Admission10140000.00%0.00%0
Social1080000.00%0.00%0
CatFailF%NP%Tot
Billing
Pend 240Skip 0Runs 2
2
0.25%
30.79%
786
Web
Pend 0Skip 2Runs 2
2
0.04%
0.09%
4668
Frontend
Pend 0Skip 0Runs 2
6
0.32%
0.32%
1902
Library
Pend 0Skip 0Runs 0
0
0.00%
0.00%
516
University
Pend 0Skip 0Runs 0
0
0.00%
0.00%
18
Subscriptions
Pend 0Skip 0Runs 0
0
0.00%
0.00%
6
Admission
Pend 0Skip 0Runs 0
0
0.00%
0.00%
1014
Social
Pend 0Skip 0Runs 0
0
0.00%
0.00%
108

Recent Runs

Recent Daily Suite Runs

DatePipelineSuitesStatusSummary
2026-07-26 18:16164728BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocialFAILEDTotal 1503 | Passed 1458 | Failed 5 | Pending 40
2026-07-27 18:16164855BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocialFAILEDTotal 1503 | Passed 1462 | Failed 1 | Pending 40
2026-07-28 18:16164959BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocialPASSEDTotal 1503 | Passed 1463 | Failed 0 | Pending 40
2026-07-29 18:15165013BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocialFAILEDTotal 1503 | Passed 1458 | Failed 3 | Pending 40
2026-07-30 18:15165111BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocialPASSEDTotal 1503 | Passed 1463 | Failed 0 | Pending 40
2026-07-31 18:16165195BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocialFAILEDTotal 1503 | Passed 1462 | Failed 1 | Pending 40
2026-07-26 18:16Pipeline 164728BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocial
FAILED
T 1503 | P 1458 | F 5 | Pend 40
2026-07-27 18:16Pipeline 164855BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocial
FAILED
T 1503 | P 1462 | F 1 | Pend 40
2026-07-28 18:16Pipeline 164959BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocial
PASSED
T 1503 | P 1463 | F 0 | Pend 40
2026-07-29 18:15Pipeline 165013BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocial
FAILED
T 1503 | P 1458 | F 3 | Pend 40
2026-07-30 18:15Pipeline 165111BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocial
PASSED
T 1503 | P 1463 | F 0 | Pend 40
2026-07-31 18:16Pipeline 165195BillingWebFrontendLibraryUniversitySubscriptionsAdmissionSocial
FAILED
T 1503 | P 1462 | F 1 | Pend 40

Recent Smoke Attempts

DateSuitePipelineJobStatusPassedFailedDuration
2026-07-27 11:58Frontend164764Frontend smokePASSED11004m 10s
2026-07-27 12:19Frontend164768Frontend smokePASSED11004m 31s
2026-07-27 12:24Frontend164771Frontend smokePASSED11004m 36s
2026-07-27 13:11Frontend164780Frontend smokePASSED11004m 10s
2026-07-27 13:46Frontend164805Frontend smokePASSED11004m 24s
2026-07-28 10:36Frontend164851Frontend smokePASSED11004m 32s
2026-07-29 17:41University165006University smokePASSED6003m 17s
2026-07-30 11:48Frontend165006Frontend smokePASSED11004m 24s
2026-07-30 13:47University165060University smokePASSED6003m 30s
2026-07-30 13:49Frontend165060Frontend smokePASSED11004m 31s
2026-07-30 15:02University165089University smokePASSED6003m 08s
2026-07-30 15:50Frontend165089Frontend smokePASSED11004m 21s
2026-07-31 11:52University165155University smokePASSED6003m 17s
2026-07-31 11:56Frontend165155Frontend smokePASSED11004m 11s
2026-07-31 12:56Frontend165165Frontend smokePASSED11004m 25s
2026-07-31 21:40University165194University smokePASSED6003m 27s
2026-07-31 22:30Frontend165194Frontend smokePASSED11004m 04s

Smoke Suite Breakdown

Frontend
12 attempts across 12 pipelines
100% green
Passed12
Failed0
Incomplete0
Avg runtime4m 22s
Median passing runtime4m 24s
Pipelines12
University
5 attempts across 5 pipelines
100% green
Passed5
Failed0
Incomplete0
Avg runtime3m 20s
Median passing runtime3m 17s
Pipelines5
Generated from GitLab project adservio/helm2. Times are shown in Europe/Bucharest. Daily-suite runtime is measured from GitLab pipeline and job timestamps. Category counts come from GitLab test-report JSON artifacts, with job-trace fallback when older artifacts have expired.