Engineering Council Test Reliability Report

Scope aligned with Slack channel #dezvoltare, covering 2026-09-12 07:00 to 2026-09-19 07:00. Metrics and timings are sourced from GitLab pipelines, jobs, and test-report artifacts for the daily 6 PM regression suite and the production smoke suite. Trend charts use daily buckets across this window.

Executive Snapshot

7
Daily Runs
0/7
Daily Green
30m 55s
Avg Daily Runtime
23
Smoke Attempts
22/23
Smoke Green
4m 34s
Avg Smoke Runtime
4m 35s
Median Smoke Time
0
Current Green Streak

Executive Analysis

Bottom line: the regression system is informative but not calm. The data suggest repeatable problem areas rather than random breakage, which means focused ownership should move the needle quickly.

What Matters

  • Daily regression passed 0 of 7 runs (0.0%), with a current green streak of 0 and a best streak of 0 in this window. The latest daily run (171445) failed, so the system is ending the week under tension rather than in a clean state. 7 failed run(s) never reached complete daily-suite counts, which points to some infrastructure or setup noise mixed into the product signal.
  • Smoke passed 22 of 23 attempts (95.7%) across 21 production pipelines. 1 pipeline(s) recovered on rerun, which is useful for continuity but also a sign that first-pass deploy signal is noisier than it should be.
  • Failure concentration is not random: Library has the highest strict failure ratio at 0.46%, while Library has the broadest non-pass footprint at 1.15%.
  • University is the weakest smoke surface in this window at 1/2 green (50.0%).
  • Daily-suite runtime averaged 30m 55s.

Engineering Analysis

  • A release gate should fail loudly for product regressions and quietly for infrastructure noise. Rerun recoveries plus incomplete daily or smoke attempts suggest those two failure modes are still partially mixed together.
  • The failure profile is concentrated enough to act on. Library and Library are carrying the strongest signal, which means reliability work should be assigned by category ownership instead of treating the suite as one undifferentiated problem.
  • The broader daily suite is carrying more instability than smoke, which usually means product regressions are escaping into wider coverage areas even when the narrow deploy gate looks acceptable.
  • The daily suite is now large enough that runtime itself is becoming a management variable at 30m 55s average duration. At that size, every additional flaky or redundant test has a measurable cost on feedback speed.

Recommended Actions

  • Split incomplete execution failures from real assertion failures in the report narrative. Setup breakage should stay visible, but it should not look identical to a product regression in the executive readout.
  • Assign one owner to Library for the next cycle and expect a short written burn-down: top failing tests, suspected root causes, flake versus regression breakdown, and what gets fixed or quarantined first.
  • Treat the daily regression suite like an operations queue until it is calm again: triage failures after each red run, close known-noise items fast, and avoid letting multiple unrelated red signals pile up between runs.
  • Put University smoke under closer guardrails for the next release cycle. It is the best place to improve first-pass deploy confidence quickly.

Improvement Ideas

  • Introduce a small reliability budget for tests: every flaky or quarantined case needs an owner and an expiry, and the team should review that budget weekly the same way it reviews bugs or incidents.
  • Track first-fail to root-cause time as a core metric. Fast diagnosis is as important as raw pass rate because the practical value of a test gate depends on how quickly it helps the team recover.
  • Define a runtime budget per suite and require justification when test count or duration grows. Reliable feedback systems stay trusted when they remain both stable and proportionate.

Category Execution Ratios

How computed

Category total executions means the sum of that category's observed test executions across every daily-suite run in the selected window.

Strict Failure Ratio = failed executions for that category divided by total executions for that category across the window.

Non-pass Ratio = (failed + pending + skipped) executions for that category divided by total executions for that category across the window.

Example: if Billing executed 800 times across the week and 2 of those executions failed, Billing strict failure ratio is 0.25%. That does not mean 0.25% of pipelines failed; it means 0.25% of observed Billing executions ended in failed.

How computed

Category total executions means the sum of that category's observed test executions across every daily-suite run in the selected window.

Strict Failure Ratio = failed executions for that category divided by total executions for that category across the window.

Non-pass Ratio = (failed + pending + skipped) executions for that category divided by total executions for that category across the window.

Example: if Billing executed 800 times across the week and 2 of those executions failed, Billing strict failure ratio is 0.25%. That does not mean 0.25% of pipelines failed; it means 0.25% of observed Billing executions ended in failed.

Daily Daily Suite Status0000109-1209-1409-1609-18
Daily Smoke Attempts0246809-1209-1409-1609-18
Daily Average Daily Suite Runtime17m 08s28m 02s38m 55s49m 48s60m 41s09-1209-1409-1609-18
Daily Average Smoke Runtime0m 00s1m 12s2m 24s3m 37s4m 49s09-1209-1409-1609-18
Daily Suite Total Test Growth (Recent 7 Runs)14716818921023209-1209-1409-1609-18
Smoke Suite Total Test Growth (Latest Run Per Day)
FrontendUniversity
60728597110Frontend 09-14: 110Frontend 09-15: 110Frontend 09-16: 110Frontend 09-17: 110Frontend 09-18: 110University 09-17: 6009-1409-1509-1609-1709-18

Category Aggregate Table

How computed

Category total executions means the sum of that category's observed test executions across every daily-suite run in the selected window.

Strict Failure Ratio = failed executions for that category divided by total executions for that category across the window.

Non-pass Ratio = (failed + pending + skipped) executions for that category divided by total executions for that category across the window.

Example: if Billing executed 800 times across the week and 2 of those executions failed, Billing strict failure ratio is 0.25%. That does not mean 0.25% of pipelines failed; it means 0.25% of observed Billing executions ended in failed.

How computed

Category total executions means the sum of that category's observed test executions across every daily-suite run in the selected window.

Strict Failure Ratio = failed executions for that category divided by total executions for that category across the window.

Non-pass Ratio = (failed + pending + skipped) executions for that category divided by total executions for that category across the window.

Example: if Billing executed 800 times across the week and 2 of those executions failed, Billing strict failure ratio is 0.25%. That does not mean 0.25% of pipelines failed; it means 0.25% of observed Billing executions ended in failed.

CategoryTotalFailedPendingSkippedFailure RatioNon-pass RatioRuns With Failures
Billing10220000.00%0.00%0
Web00000.00%0.00%7
Frontend00000.00%0.00%7
Library4332030.46%1.15%2
CatFailF%NP%Tot
Billing
Pend 0Skip 0Runs 0
0
0.00%
0.00%
1022
Web
Pend 0Skip 0Runs 7
0
0.00%
0.00%
0
Frontend
Pend 0Skip 0Runs 7
0
0.00%
0.00%
0
Library
Pend 0Skip 3Runs 2
2
0.46%
1.15%
433

Recent Runs

Recent Daily Suite Runs

DatePipelineSuitesStatusSummary
2026-09-12 18:24170442BillingWebFrontendLibraryFAILEDTotal 232 | Passed 232 | Failed 0 | Incomplete suite counts
2026-09-13 18:21170444BillingWebFrontendLibraryFAILEDTotal 232 | Passed 232 | Failed 0 | Incomplete suite counts
2026-09-14 18:22170690BillingWebFrontendLibraryFAILEDTotal 147 | Passed 146 | Failed 1 | Incomplete suite counts
2026-09-15 19:03170793BillingWebFrontendLibraryFAILEDTotal 232 | Passed 232 | Failed 0 | Incomplete suite counts
2026-09-16 18:20170903BillingWebFrontendLibraryFAILEDTotal 148 | Passed 148 | Failed 0 | Incomplete suite counts
2026-09-17 19:03171125BillingWebFrontendLibraryFAILEDTotal 232 | Passed 232 | Failed 0 | Incomplete suite counts
2026-09-18 18:22171445BillingWebFrontendLibraryFAILEDTotal 232 | Passed 228 | Failed 1 | Incomplete suite counts
2026-09-12 18:24Pipeline 170442BillingWebFrontendLibrary
FAILED
T 232 | P 232 | F 0 | Pend 0 | Incomplete
2026-09-13 18:21Pipeline 170444BillingWebFrontendLibrary
FAILED
T 232 | P 232 | F 0 | Pend 0 | Incomplete
2026-09-14 18:22Pipeline 170690BillingWebFrontendLibrary
FAILED
T 147 | P 146 | F 1 | Pend 0 | Incomplete
2026-09-15 19:03Pipeline 170793BillingWebFrontendLibrary
FAILED
T 232 | P 232 | F 0 | Pend 0 | Incomplete
2026-09-16 18:20Pipeline 170903BillingWebFrontendLibrary
FAILED
T 148 | P 148 | F 0 | Pend 0 | Incomplete
2026-09-17 19:03Pipeline 171125BillingWebFrontendLibrary
FAILED
T 232 | P 232 | F 0 | Pend 0 | Incomplete
2026-09-18 18:22Pipeline 171445BillingWebFrontendLibrary
FAILED
T 232 | P 228 | F 1 | Pend 0 | Incomplete

Recent Smoke Attempts

DateSuitePipelineJobStatusPassedFailedDuration
2026-09-14 01:37Frontend170447Frontend smokePASSED11004m 58s
2026-09-14 21:49Frontend170695Frontend smokePASSED11004m 24s
2026-09-15 09:13Frontend170703Frontend smokePASSED11004m 30s
2026-09-15 09:28Frontend170706Frontend smokePASSED11004m 25s
2026-09-15 09:58Frontend170711Frontend smokePASSED11004m 20s
2026-09-15 10:13Frontend170717Frontend smokePASSED11004m 35s
2026-09-15 11:30Frontend170736Frontend smokePASSED11004m 23s
2026-09-15 12:38Frontend170742Frontend smokePASSED11004m 23s
2026-09-15 15:51Frontend170777Frontend smokePASSED11004m 40s
2026-09-16 02:02Frontend170819Frontend smokePASSED11005m 08s
2026-09-16 13:41Frontend170879Frontend smokePASSED11004m 31s
2026-09-16 18:01Frontend170902Frontend smokePASSED11004m 48s
2026-09-16 21:06Frontend170909Frontend smokePASSED11004m 35s
2026-09-17 11:45Frontend170960Frontend smokePASSED11004m 28s
2026-09-17 11:58University170960University smokeFAILED5913m 52s
2026-09-17 13:03Frontend171007Frontend smokePASSED11004m 41s
2026-09-17 16:01University171115University smokePASSED6004m 03s
2026-09-17 16:04Frontend171115Frontend smokePASSED11004m 37s
2026-09-17 17:22Frontend171121Frontend smokePASSED11004m 36s
2026-09-17 20:52Frontend171137Frontend smokePASSED11004m 52s
2026-09-17 22:01Frontend171141Frontend smokePASSED11004m 26s
2026-09-18 18:33Frontend171436Frontend smokePASSED11004m 56s
2026-09-18 19:08Frontend171459Frontend smokePASSED11004m 42s

Smoke Suite Breakdown

Frontend
21 attempts across 21 pipelines
100% green
Passed21
Failed0
Incomplete0
Avg runtime4m 37s
Median passing runtime4m 35s
Pipelines21
University
2 attempts across 2 pipelines
50% green
Passed1
Failed1
Incomplete0
Avg runtime3m 58s
Median passing runtime4m 03s
Pipelines2
Generated from GitLab project adservio/helm2. Times are shown in Europe/Bucharest. Daily-suite runtime is measured from GitLab pipeline and job timestamps. Category counts come from GitLab test-report JSON artifacts, with job-trace fallback when older artifacts have expired.