This page gets a project under version control and connected to github, and sketches the two ways you will actually touch course notebooks in Colab. It assumes nothing: install git once, and the loop below — edit, stage, commit — is the same whether the project is a five-line script or a full thesis repository.
Why version control matters for a large, collaborative project¶
Any project reworked over time, shared with collaborators, or built on data that keeps changing outgrows what a folder of files can track. Without version control, that history usually survives as files named analysis.py, analysis_v2.py, analysis_final.py — a manual, error-prone imitation of what git does for free. A thesis is one example of a project like this; so is a group assignment or a shared analysis pipeline.
git replaces the filename as a version marker with an actual, searchable history: every commit records who changed what, when, and why, and any of them can be recovered exactly. That matters in practice: you can always get back to a version a collaborator last reviewed; when someone asks how a figure or a number was produced, you can point at the exact commit that produced it; and the project’s code, methodology, and results accumulate a paper trail instead of a pile of loose files. github then gives that history a home outside any one laptop — a backup, and a place to share the repository with the rest of the team.
Installing and configuring git¶
Check whether git is already installed with git --version. If not: macOS ships it with the Xcode Command Line Tools (xcode-select --install), or install it with brew install git; Linux users typically already have it, or can get it with their package manager (sudo apt install git on Debian/Ubuntu); Windows users install Git for Windows, which also provides a unix-like terminal.
Then set your identity once — it is attached to every commit you make:
git config --global user.name "Your Name"
git config --global user.email "you@example.com"The core git loop: init, status, add, commit, log¶
git records snapshots of a project so you can review history, revert mistakes, and collaborate. The core loop is: edit files, stage the changes you want to keep, then commit them with a message. git status shows the current state at every step.
git init # start tracking the current folder
git status # what has changed, what is staged
git add analysis.py # stage one file (git add . stages all)
git commit -m "Add lapse-rate fit" # record a snapshot with a message
git log --oneline # compact historyA commit is a labelled checkpoint. Commit small, coherent changes often; a good message says why, not just what.
Branches: isolating work in progress¶
Branches let you develop a change in isolation, then merge it back.
git switch -c experiment # create and move onto a new branch (older git: git checkout -b)
# ... edit and commit on 'experiment' ...
git switch main # return to the main branch
git merge experiment # bring the experiment's commits into maingithub: account, repository, and remotes¶
Create a free account at github.com, then a repository — either starting from github, or starting locally and connecting it afterward.
Starting from github gives you a git clone you build on directly:
git clone git@github.com:your-user/your-repo.git
cd your-repoStarting from an existing local project, create an empty repository on github (no README, no .gitignore — those would conflict with your existing files), then point your local project at it as a remote:
git remote add origin git@github.com:your-user/your-repo.git
git push -u origin main # first push sets the upstream; later just `git push`
git pull # fetch and merge remote changesBoth commands above use ssh (the git@github.com:... form) rather than https, which needs a keypair set up first — see the dedicated SSH page.
Colab: two ways to work with course notebooks¶
Every notebook in this book has an “Open in Colab” link. Which way you use it depends on whether you plan to keep your own changes.
Quick, read-only access. Click the badge; Colab opens a live copy of the notebook straight from the course github repository. You can run and edit cells freely, but nothing is saved back to github — Colab keeps changes only in your own Google Drive, if at all, and opening the badge again gives you the clean original. This is enough for working through a subchapter’s cells or exercises.
Your own tracked copy. Clone the course repository (or your own fork of it) as above, then either open notebooks locally in Jupyter, or mount your Drive in Colab and open the notebook from your cloned folder. Now your edits are files in a git repository you control: commit them, push them, and they persist and version like any other project. This is the route to take once you are doing your own analysis rather than following along — including, eventually, your thesis code.
How the github/Colab integration works¶
Colab reads a notebook straight from a github URL by swapping the domain: a notebook at
https://github.com/your-org/your-repo/blob/main/notebook.ipynbopens in Colab at
https://colab.research.google.com/github/your-org/your-repo/blob/main/notebook.ipynbThat is exactly what an “Open in Colab” badge is — a link built from this pattern, pointing at one specific notebook on one specific branch. The integration runs both directions:
github -> Colab (reading). Opening that link, or using Colab’s File -> Open notebook -> GitHub tab, loads the notebook’s current content straight from github. Nothing is cloned or downloaded to your machine; Colab reads the file directly.
Colab -> github (writing). File -> Save a copy in GitHub commits your edits back. The first time, Colab asks permission to access your github account; after that, it lets you pick the repository, branch, and a commit message, and pushes a real commit — no terminal involved. This is a genuine alternative to the clone-then-git push route from the section above, useful when you want to edit a notebook entirely inside Colab.
Going deeper: .gitignore
A .gitignore file lists patterns git should not track — keep data, environments, and caches out of history.
# environments and caches
.venv/
__pycache__/
*.pyc
# data and outputs (version the code, archive the data elsewhere)
data/
*.nc
*.zarr/Commit the .gitignore itself. Large or binary data does not belong in git; archive it with a doi instead (see the Zenodo box below).
Going deeper: pull requests
A pull request proposes merging one branch into another, and gives your advisor or collaborators a place to comment before it lands.
git switch -c fix-lapse-rate
# ... edit and commit ...
git push -u origin fix-lapse-rate
# open a pull request on github.com comparing it to main, review, then merge thereThis is the standard way to combine everyone’s changes to the same repository without commits landing on main unreviewed.
Going deeper: archiving code and data with a Zenodo doi
Zenodo mints a permanent, citable doi for a github release. Link your github account to Zenodo, enable the repository, then publish a tagged release on github; Zenodo archives that snapshot and issues a doi. Cite the doi in papers so others can retrieve the exact version used. This is the standard route to making code and small datasets reproducible and FAIR.
Resources¶
Software Carpentry — Version Control with Git — the standard first course on the git workflow used above.
Project Pythia — Getting started with GitHub — geoscience-oriented walkthrough of github, ssh setup, cloning, and branches.
GitHub Docs — Hello World — github’s own quickstart: repository, branch, commit, pull request.