Skip to content
ImpactMojo ImpactMojo
Browse Membership
Back to Code Studio
Code Studio · Guided tool course

Git and Quarto: reproducible reports

Keep a record of every change to your analysis, share it with colleagues without emailing files, and turn one report template into ten district profiles with a single command. Git records versions; Quarto turns R or Python code and text into HTML, Word or PDF reports. Both are free. Cells on this page run the parts that can run in a browser.

Git and Quarto run on your own computer, in a terminal. The grey boxes are commands and files for you to type there; this page shows no terminal output from them, because it cannot run them. Each module says what to look for. The code cells run the steps a browser can: a diff, a check for personal data before a commit, and the calculation inside a district profile. The first run downloads the engine once (R about 7 MB, Python about 10 MB).
Module 1 of 8

Why version control, and installing Git

A folder of report_final.docx, report_final_v2.docx and report_final_v2_AK_comments.docx is a version control system with no record of what changed, who changed it or why. Git keeps that record. Each commit is a saved snapshot of the project with a message. You can see the difference between any two snapshots, go back to one, and work on a change without disturbing the version others rely on.

Git suits text files: R and Python scripts, Quarto documents, CSV codebooks, XLSForms saved as CSV. It stores Word and Excel files but cannot show what changed inside them.

Version and cost, checked 6 October 2026. The Git home page lists the latest source release as 2.56.0 (release notes dated 28 September 2026). Git is "released under the GNU General Public License version 2.0", so it costs nothing. The Windows download page offers Git for Windows 2.56.0(2), released 5 October 2026.

Install

First-time setup

These lines come from the Pro Git book's setup chapter. Your name and email are written into every commit you make.

Terminal: set up Git once
# Tell Git who you are (once per computer). Use the name and email you want on your commits.
git config --global user.name "Asha Kumari"
git config --global user.email "asha@example.org"

# Name the first branch of every new repository "main"
git config --global init.defaultBranch main

# Check what Git has recorded
git config --list
git --version
Exercise. Run git --version after installing. It should print a version number starting with git version. If the terminal says the command is not found, close it, open a new one, and try again; on Windows, use Git Bash.
Module 2 of 8

The basic loop: init, add, commit, log, diff

Almost all daily Git work is four commands: change files, git add the ones you want in the snapshot, git commit with a message, and look back with git log and git diff.

Terminal: the basic loop
cd district-profiles          # your project folder
git init                      # start a repository here (creates a hidden .git folder)
git status                    # what has changed since the last commit?

git add README.md analysis/profile.qmd     # stage these files for the next commit
git commit -m "Add district profile template"

# ...edit analysis/profile.qmd...
git diff                      # line-by-line changes not yet staged
git add analysis/profile.qmd
git commit -m "Report toilet coverage as a percentage"

git log --oneline             # one line per commit, newest first

Reading a diff

Git shows changes in the unified diff format: a line starting with - was removed and a line starting with + was added, with unchanged lines around them for context. Python's difflib writes the same format, so the cell shows what git diff would report after a cleaning fix changed one figure in a district table (figures from households.csv, illustrative data invented for teaching).

The output names the old and new file, then shows the Purnia row twice: once with - and 75.0, once with + and 79.2. Barmer and Rewa appear unmarked, as context. This is why Git suits CSV tables and code: a reviewer sees exactly which number moved.

Exercise. In the after text, also change Rewa's mean spending from 2557 to 2575, and run again. Both changes appear. Then add a new line for Gaya at the end of after: it appears with a + and no matching -.
Module 3 of 8

Branches: try a change without breaking main

A branch is a separate line of commits. Keep main as the version that works, and make a branch for anything that might not: a new table, a different poverty line, a colleague's suggested rewrite.

Terminal: branches
git branch                         # list branches; * marks the one you are on
git switch -c caste-tables         # create a branch and move to it
# ...edit, add, commit on the branch...
git switch main                    # back to main; the branch's work is set aside
git merge caste-tables             # bring the branch's commits into main
git branch -d caste-tables         # delete the branch once merged
Exercise. In a practice folder, commit a file with one line, create a branch, change the line and commit. Switch back to main and open the file: it shows the old line. Merge the branch and open it again. Run git log --oneline at each step to see where you are.
Module 4 of 8

Remotes: GitHub, push and pull

A remote is a copy of the repository on a server, which is how a team shares work and how a laptop's work survives the laptop. GitHub is the most common host. Its documentation says a free account can own unlimited public and private repositories, and that private repositories are only accessible to you and the people you share them with.

  1. On GitHub, create a new repository. Choose Private. Leave it empty (no README) if your folder already has commits.
  2. Copy its HTTPS address and add it as a remote called origin, as below.
  3. Push. GitHub's remote repositories guide says that when Git asks for your password over HTTPS, you enter a personal access token, because password authentication for Git has been removed. The same guide mentions Git Credential Manager as an alternative that remembers your login.
Terminal: remotes
# Connect the folder to an empty private repository you created on GitHub
git remote add origin https://github.com/YOUR-ORG/district-profiles.git
git remote -v                       # check the address
git push -u origin main             # send main to GitHub; -u remembers the pairing

# Every working day
git pull                            # get colleagues' commits first
# ...work, add, commit...
git push                            # send your commits

# A colleague starting on a new laptop
git clone https://github.com/YOUR-ORG/district-profiles.git
Pull before you push. If a colleague pushed first, your push is refused until you pull their commits. Pull at the start of each day and before each push, and conflicts stay small.
Exercise. Create a private repository, push your practice folder, then edit README.md on the GitHub website and commit it there. Back on your computer, run git pull and open README.md: the website edit is now in your folder.
Module 5 of 8

.gitignore and keeping personal data out

Every file you commit stays in the repository's history, on every computer that clones it and on the server. A beneficiary list with names and phone numbers committed once is, for practical purposes, permanent. Under the DPDP Act 2023, section 8(5), which applies from 13 May 2027, your organisation must take "reasonable security safeguards to prevent personal data breach". A private repository is access control, and it holds only as long as the access list stays right. Keep personal data out of Git from the start.

A .gitignore file lists files Git should not track. From the gitignore manual: a line starting with # is a comment; a pattern ending in / matches only directories; * "matches anything except a slash".

.gitignore
# .gitignore at the top of the repository

# Personal data never goes into the repository
data/raw/
*.xlsx
*.sav
*.dta

# Rendered output can be rebuilt from the code
outputs/
*_files/

# Credentials and local settings
.env
.Rhistory
.Rproj.user/

The cell builds a small practice project with an invented beneficiary list, applies a simplified version of these rules, and scans every file for strings that look like Indian mobile numbers (ten digits starting with 6 to 9).

Three files would be committed (README.md, the .qmd and the clean household file) and three are ignored: the raw beneficiary list in data/raw/, the Excel file, and the rendered output. The scan finds 2 phone-like numbers, both in the ignored raw file.

Exercise. Delete the data/raw/ line from gitignore and run again. The beneficiary list is now marked COMMIT with its two numbers flagged. A check like this, run before every commit, catches the mistake while it is still on your laptop. The cell's matcher is simpler than Git's; on your computer, git status --ignored lists what Git itself ignores.

If personal data was already committed

The gitignore manual says files already tracked are not affected by .gitignore, and points to git rm --cached, which removes a file from the index while leaving it on disk (git rm manual).

Terminal: stop tracking a file
# A file was committed before it was listed in .gitignore.
# Stop tracking it (the file stays on your disk), then commit.
git rm --cached data/raw/beneficiaries.csv
git commit -m "Stop tracking raw beneficiary list"
# The file is still inside every earlier commit. If it was ever pushed,
# treat it as disclosed: tell your data protection lead.
Removing a file in a new commit does not remove it from history. Anyone with a clone, or with access to the remote, can still check out the earlier commit. If the file was pushed, report it to whoever handles data protection in your organisation, as a possible breach, and get help to rewrite the history. From 13 May 2027, section 8(6) of the DPDP Act requires a Data Fiduciary to inform the Data Protection Board and each affected person of a personal data breach.
Module 6 of 8

Quarto documents: text and code in one file

A Quarto document (.qmd) is a plain text file: a YAML header between two --- lines, then text in Markdown and code chunks in R or Python. Rendering runs the code and writes the results into the report, so the numbers in the text always come from the data.

Version and cost, checked 6 October 2026. The Quarto download page's current release is 1.10.19, published 6 October 2026. The licence page says the Quarto command-line tool "is licensed under the MIT License (version 1.4 and later)". It is free. Installers are listed for Windows, macOS and Linux.
profile.qmd (R)
---
title: "District profile"
author: "MEL team"
date: today
format:
  html:
    embed-resources: true
  docx: default
execute:
  echo: false
---

## Coverage

```{r}
#| label: load
hh <- read.csv("households.csv")
```

Households surveyed: `r nrow(hh)`.

```{r}
#| label: toilet-by-district
#| warning: false
round(100 * tapply(hh$has_toilet == "Yes", hh$district, mean), 1)
```
profile.qmd (Python)
---
title: "District profile"
format: html
jupyter: python3
---

```{python}
#| echo: false
import pandas as pd
hh = pd.read_csv("households.csv")
print(len(hh), "households")
```

Render to HTML, Word and PDF

Terminal: render
quarto preview profile.qmd               # render and open in a browser, re-rendering on save
quarto render profile.qmd --to html      # a single HTML file
quarto render profile.qmd --to docx      # a Word document
quarto render profile.qmd --to pdf       # a PDF (needs a TeX installation, below)

quarto install tinytex                   # the TeX distribution Quarto recommends for PDF

The Quarto tutorial notes that the file name should come first after quarto render. For PDF, the PDF guide says you need a TeX distribution and recommends TinyTeX, installed with the last command above. Word output needs nothing extra.

Exercise. Save the R version as profile.qmd in a folder with households.csv, run quarto render profile.qmd --to html, and open the HTML file. Look for the sentence "Households surveyed: 240." and a table of ten district percentages. Then render --to docx and open the Word file: the same numbers, no copying.
Module 7 of 8

Parameterised reports: one template, ten district profiles

A parameter is a value the report reads from outside, such as the district name. Write the report once, then render it once per district. The Quarto parameters guide gives one syntax for each engine.

profile.qmd with a parameter (R)
---
title: "District profile: `r params$district`"
format: html
params:
  district: "Purnia"
---

```{r}
hh <- read.csv("households.csv")
g  <- hh[hh$district == params$district, ]
```

`r params$district` has `r nrow(g)` surveyed households.
The parameters cell (Python)
```{python}
#| tags: [parameters]
district = "Purnia"
```
Terminal: render with parameters
# One district
quarto render profile.qmd -P district:Gaya --output gaya.html

# Every district, in a terminal (macOS, Linux, or Git Bash on Windows)
for d in Barmer Betul Gaya Indore Kozhikode Patna Purnia Rewa Udaipur Wayanad; do
  quarto render profile.qmd -P district:$d --output profile-$d.html
done

-P district:Gaya sets the parameter (the guide's own example is -P alpha:0.2) and --output names the file, so the ten renders do not overwrite each other.

The cell runs what the report's chunk computes, for one district. Switch between R and Python with the tabs on the cell.

For Purnia (Bihar) it reports 24 households, 75.0% with a toilet, 75.0% with a bank account and mean monthly spending of Rs 2,395 per person, then the count of households in each caste group. Your rendered Purnia profile should show the same figures.

Exercise. Change the district to Kozhikode and run again; then to Kozhikkode. The misspelt name finds no households: R prints 0 households and NaN percentages, and Python stops with an error when it looks up the state. Add a check at the top of your .qmd that stops the render when the district is not in districts.csv, so a typo in a loop cannot produce an empty profile with a confident title.
Module 8 of 8

A reproducible project layout

A project is reproducible when someone else can take the repository and the raw data, run it, and get the same report. A plain folder layout and three rules get most of the way.

Project folders
district-profiles/
  README.md              what this is, how to run it, who to ask
  .gitignore             data/raw/, outputs/, credentials
  data/
    raw/                 exports exactly as downloaded: never edited, never committed
    clean/               written only by the cleaning script
  scripts/
    01_clean.R           raw -> clean, with checks that stop on a problem
  analysis/
    profile.qmd          the parameterised report
  outputs/               rendered reports: rebuilt, never edited by hand
Exercise. Set up this layout for a real project, with the .gitignore from module 5. Commit, push to a private repository, then clone it into a new folder, copy the raw data in, and run the README's command. If the reports come out the same, the project is reproducible. If something is missing, the clone tells you what.

Where next

→

R and Python side by side

The language for the code chunks in your reports.

→

The tidyverse

Write the tables and charts your Quarto profiles will hold.

→

Spreadsheets for M&E

Where most of the numbers start, and when to leave them.

→

Data Protection and the DPDP Act

What the law asks before personal data goes anywhere.

→

Research Ethics 101

Consent, confidentiality and data handling in field research.