Developing with ChatGPT

Use ChatGPT to analyze CI metrics you can reproduce

ChatGPT

Inspect the input, define the comparison, and return calculations and a report that another developer can check.

Applies to
ChatGPT
Last verified
Reviewed by
Timothy Fehr

A build dashboard says the average duration is unchanged, but developers say the pipeline feels slower. Use ChatGPT to inspect the underlying records and prepare a reproducible analysis.

ChatGPT Work can work with data files and produce analysis artifacts. Check that your session has the file and execution tools the task needs. If it cannot execute calculations, request a script and run it yourself.

Start with a small dataset

This synthetic CSV contains six runs of the same job:

run_id,period,job,duration_seconds,status
b1,baseline,test,100,passed
b2,baseline,test,100,passed
b3,baseline,test,400,passed
c1,candidate,test,200,passed
c2,candidate,test,200,passed
c3,candidate,test,200,passed

Save it as ci-runs.csv or attach equivalent approved data. Give units and define which runs belong in each period. For real exports, also account for runner type, cache state, cancellation, retries, and duplicate records.

Compare baseline and candidate durations for the test job. Inspect types, missing values, duplicate run IDs, and status values before calculating. Show sample counts, mean, median, and maximum. Keep cancelled or failed runs separate if they appear. Return the analysis code and a short report with its assumptions. Do not claim the change caused a difference from this sample.

Inspect before interpreting

The example has three records per period and no missing durations. A useful analysis shows those counts before discussing the trend. If an export includes two jobs with the same run ID, the uniqueness key needs both run and job.

Keep exclusions visible in a separate table. Dropping failed builds can make a slow or unreliable pipeline look fast. Changing job mix can also change the aggregate without any individual job becoming slower.

Ask for both a compact result table and the code that produced it:

PeriodRunsMean secondsMedian secondsMaximum seconds
Baseline3200100400
Candidate3200200200

The means match, while the median doubles and the maximum falls. With three runs per period, this is a description of the example, not evidence of a stable production trend.

Return the work, not just a chart

The useful deliverable includes the input definition, cleaning decisions, analysis code, result table, and the unanswered engineering question. A chart should show its units and periods. Preserve the raw CSV so a reviewer can rerun the analysis after questioning an exclusion.

A next investigation might compare matched runner types over more runs. Choose that follow-up from the unresolved evidence rather than asking for a more confident explanation of the same six records.

What goes wrong

A duration column can be read as text. Duplicate exports can count a run twice. The model may compare milliseconds with seconds, hide exclusions, or explain a causal relationship the data does not establish.

Request the inspection output and executable calculation. A polished report without those artifacts leaves the most important assumptions invisible.

How to check

Count the six records. Each period sums to 600 seconds, giving a mean of 200. Sort the three values in each period to check the medians. Confirm the maxima directly from the CSV.

Run the supplied code against the original input. Add a duplicate record and a blank duration to a copy and confirm they are surfaced explicitly. If the tool could not execute, label that limitation in the report before sharing it.

Sources

  1. OpenAI: analyze datasets and ship reports Tier 1 2026-09-08