Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

AI chatbots produce plausible code fast, but cannot know whether that code is right for your data, your units, or your question. This page is a living guide — expect it to change faster than the rest of the book as models and norms evolve — built around three habits: being a precise communicator, a skeptical reviewer, and an active learner.

Be a precise communicator

A prompt that gets a useful answer on the first try usually states four things:

  1. The goal — the high-level scientific objective, not just the mechanical step.

  2. The environment — language and libraries (python, pandas, numpy, ...).

  3. The data — its actual structure; paste a .head() or print() output rather than describing it from memory.

  4. The task — the precise action wanted.

Compare two prompts for the same task: averaging a temperature record by month.

A bad prompt:

How do I average my data by month in python?

Without a description of the data, the assistant reaches for a generic example — a three-row toy table that happens to already have monthly, not daily, dates — and returns code that runs but has nothing to do with the actual station data.

A good prompt:

I am using python with the pandas library to analyze weather station data. I have a DataFrame df_temp with a DatetimeIndex and a column air_temp_celsius. Here is df_temp.head():

                      air_temp_celsius
2024-07-01 00:00:00              15.2
2024-07-01 01:00:00              14.9

Please give me code to compute the mean monthly air temperature, stored in a new DataFrame monthly_mean_temp.

Pinning the actual column names and structure gets df_temp.resample("ME").mean() — the right tool, applied to the right data — on the first try.

Be a skeptical reviewer

Generated code runs and looks reasonable far more often than it is actually correct — the failures that matter in science are silent, not crashes. Three checks catch most of them:

  1. Understand. Can you explain what every line does? If not, ask for a line-by-line explanation before running it.

  2. Test. Run it on a small, known subset first, and check one value by hand.

  3. Question. Is this the best approach, or just an approach? Ask for alternatives and trade-offs.

A first request — “plot a 30-day rolling average of temperature” — gets a working but bare plot. Two follow-ups sharpen it without starting over. First:

This code works, but the plot is not very readable. How would you modify it so it’s more readable?

adds axis labels, a legend, and the raw data as context. Then:

For climatological analysis, a centered mean is usually more appropriate. How would you modify this to use a centered 30-day window?

is the kind of correction only a reviewer who understands the domain, not just the syntax, would think to ask for — rolling(window=30, center=True) in place of the assistant’s default trailing window.

Be an active learner

Three prompt patterns turn the assistant into a tutor instead of a code vending machine:

A concrete failure: the unstated unit

Asked to flag freezing conditions, an assistant might write:

def is_freezing(temperature):     # unit unspecified: the latent bug
    return temperature < 0

Called on a station that reports temperature in kelvin, this is wrong for every real value: is_freezing(268.15) — a genuine −5 °C — returns False, because nothing is ever below zero kelvin. The code is not wrong in isolation; it is wrong because the prompt never fixed the unit, so the assistant guessed celsius. The fix states the unit everywhere it can — the name, the type hint, the threshold:

def is_freezing_kelvin(temp_kelvin: float) -> bool:
    return temp_kelvin < 273.15

This page will change

Models, tools, and institutional norms around AI assistance are all moving faster than the rest of this book. Treat the three habits above — precise, skeptical, active — as the stable part; the specific tools and examples will be revisited as they change.