Git and Quarto: reproducible reports
Keep a record of every change to your analysis, share it with colleagues without emailing files, and turn one report template into ten district profiles with a single command. Git records versions; Quarto turns R or Python code and text into HTML, Word or PDF reports. Both are free. Cells on this page run the parts that can run in a browser.
Why version control, and installing Git
A folder of report_final.docx, report_final_v2.docx and report_final_v2_AK_comments.docx is a version control system with no record of what changed, who changed it or why. Git keeps that record. Each commit is a saved snapshot of the project with a message. You can see the difference between any two snapshots, go back to one, and work on a change without disturbing the version others rely on.
Git suits text files: R and Python scripts, Quarto documents, CSV codebooks, XLSForms saved as CSV. It stores Word and Excel files but cannot show what changed inside them.
Install
- Windows: download the installer from the Windows page. The default choices are fine. It adds Git Bash, a terminal where every command on this page works. A portable version is offered for computers where you cannot install software.
- macOS: the macOS page says Apple ships Git with the Xcode Command Line Tools: run
xcode-select --installin Terminal. With Homebrew,brew install git. - Linux: Debian and Ubuntu,
sudo apt-get install git; Fedora,sudo dnf install git(Linux page).
First-time setup
These lines come from the Pro Git book's setup chapter. Your name and email are written into every commit you make.
# Tell Git who you are (once per computer). Use the name and email you want on your commits. git config --global user.name "Asha Kumari" git config --global user.email "asha@example.org" # Name the first branch of every new repository "main" git config --global init.defaultBranch main # Check what Git has recorded git config --list git --version
git --version after installing. It should print a version number starting with git version. If the terminal says the command is not found, close it, open a new one, and try again; on Windows, use Git Bash.The basic loop: init, add, commit, log, diff
Almost all daily Git work is four commands: change files, git add the ones you want in the snapshot, git commit with a message, and look back with git log and git diff.
cd district-profiles # your project folder git init # start a repository here (creates a hidden .git folder) git status # what has changed since the last commit? git add README.md analysis/profile.qmd # stage these files for the next commit git commit -m "Add district profile template" # ...edit analysis/profile.qmd... git diff # line-by-line changes not yet staged git add analysis/profile.qmd git commit -m "Report toilet coverage as a percentage" git log --oneline # one line per commit, newest first
git initcreates an empty repository: the manual describes it as "basically a .git directory". Run it once per project.git statusis the command to run whenever you are unsure. It lists files that changed, files staged for the next commit, and files Git is not tracking.git addstages a file. Only staged changes go into the next commit, so you can commit one fix at a time.git commit -msaves the snapshot. Write the message as what the commit does: "Fix Purnia district code", "Add caste table".git diffshows changed lines.git log --onelinelists commits with a short ID you can refer to.
Reading a diff
Git shows changes in the unified diff format: a line starting with - was removed and a line starting with + was added, with unchanged lines around them for context. Python's difflib writes the same format, so the cell shows what git diff would report after a cleaning fix changed one figure in a district table (figures from households.csv, illustrative data invented for teaching).
The output names the old and new file, then shows the Purnia row twice: once with - and 75.0, once with + and 79.2. Barmer and Rewa appear unmarked, as context. This is why Git suits CSV tables and code: a reviewer sees exactly which number moved.
after text, also change Rewa's mean spending from 2557 to 2575, and run again. Both changes appear. Then add a new line for Gaya at the end of after: it appears with a + and no matching -.Branches: try a change without breaking main
A branch is a separate line of commits. Keep main as the version that works, and make a branch for anything that might not: a new table, a different poverty line, a colleague's suggested rewrite.
git branch # list branches; * marks the one you are on git switch -c caste-tables # create a branch and move to it # ...edit, add, commit on the branch... git switch main # back to main; the branch's work is set aside git merge caste-tables # bring the branch's commits into main git branch -d caste-tables # delete the branch once merged
git switch -c namecreates the branch and moves to it; the git switch manual describes-cas "Create a new branch" before switching to it. Older tutorials usegit checkout -bfor the same thing.- Files in your folder change when you switch branches. Commit before switching.
git mergebrings the branch's commits into the branch you are on. If both changed the same lines, Git stops and marks the conflict in the file for you to resolve, then you add and commit.
main and open the file: it shows the old line. Merge the branch and open it again. Run git log --oneline at each step to see where you are.Remotes: GitHub, push and pull
A remote is a copy of the repository on a server, which is how a team shares work and how a laptop's work survives the laptop. GitHub is the most common host. Its documentation says a free account can own unlimited public and private repositories, and that private repositories are only accessible to you and the people you share them with.
- On GitHub, create a new repository. Choose Private. Leave it empty (no README) if your folder already has commits.
- Copy its HTTPS address and add it as a remote called
origin, as below. - Push. GitHub's remote repositories guide says that when Git asks for your password over HTTPS, you enter a personal access token, because password authentication for Git has been removed. The same guide mentions Git Credential Manager as an alternative that remembers your login.
# Connect the folder to an empty private repository you created on GitHub git remote add origin https://github.com/YOUR-ORG/district-profiles.git git remote -v # check the address git push -u origin main # send main to GitHub; -u remembers the pairing # Every working day git pull # get colleagues' commits first # ...work, add, commit... git push # send your commits # A colleague starting on a new laptop git clone https://github.com/YOUR-ORG/district-profiles.git
git pull and open README.md: the website edit is now in your folder..gitignore and keeping personal data out
Every file you commit stays in the repository's history, on every computer that clones it and on the server. A beneficiary list with names and phone numbers committed once is, for practical purposes, permanent. Under the DPDP Act 2023, section 8(5), which applies from 13 May 2027, your organisation must take "reasonable security safeguards to prevent personal data breach". A private repository is access control, and it holds only as long as the access list stays right. Keep personal data out of Git from the start.
A .gitignore file lists files Git should not track. From the gitignore manual: a line starting with # is a comment; a pattern ending in / matches only directories; * "matches anything except a slash".
# .gitignore at the top of the repository # Personal data never goes into the repository data/raw/ *.xlsx *.sav *.dta # Rendered output can be rebuilt from the code outputs/ *_files/ # Credentials and local settings .env .Rhistory .Rproj.user/
The cell builds a small practice project with an invented beneficiary list, applies a simplified version of these rules, and scans every file for strings that look like Indian mobile numbers (ten digits starting with 6 to 9).
Three files would be committed (README.md, the .qmd and the clean household file) and three are ignored: the raw beneficiary list in data/raw/, the Excel file, and the rendered output. The scan finds 2 phone-like numbers, both in the ignored raw file.
data/raw/ line from gitignore and run again. The beneficiary list is now marked COMMIT with its two numbers flagged. A check like this, run before every commit, catches the mistake while it is still on your laptop. The cell's matcher is simpler than Git's; on your computer, git status --ignored lists what Git itself ignores.If personal data was already committed
The gitignore manual says files already tracked are not affected by .gitignore, and points to git rm --cached, which removes a file from the index while leaving it on disk (git rm manual).
# A file was committed before it was listed in .gitignore. # Stop tracking it (the file stays on your disk), then commit. git rm --cached data/raw/beneficiaries.csv git commit -m "Stop tracking raw beneficiary list" # The file is still inside every earlier commit. If it was ever pushed, # treat it as disclosed: tell your data protection lead.
Quarto documents: text and code in one file
A Quarto document (.qmd) is a plain text file: a YAML header between two --- lines, then text in Markdown and code chunks in R or Python. Rendering runs the code and writes the results into the report, so the numbers in the text always come from the data.
- With R: install the rmarkdown package,
install.packages("rmarkdown"). The Quarto R guide says this also installs knitr, which runs R chunks. - With Python: install Jupyter,
python3 -m pip install jupyter(on Windows,py -m pip install jupyter), from the Quarto Python guide.
---
title: "District profile"
author: "MEL team"
date: today
format:
html:
embed-resources: true
docx: default
execute:
echo: false
---
## Coverage
```{r}
#| label: load
hh <- read.csv("households.csv")
```
Households surveyed: `r nrow(hh)`.
```{r}
#| label: toilet-by-district
#| warning: false
round(100 * tapply(hh$has_toilet == "Yes", hh$district, mean), 1)
```
- The YAML header sets the title and the output formats.
embed-resources: truemakes a single HTML file with no separate folder, which is easier to email. - Lines starting with
#|at the top of a chunk are chunk options:label,echo: false(hide the code, show the result),warning: false. Underexecute:in the header they apply to every chunk. - Inline R,
`r nrow(hh)`, puts a number into a sentence.
---
title: "District profile"
format: html
jupyter: python3
---
```{python}
#| echo: false
import pandas as pd
hh = pd.read_csv("households.csv")
print(len(hh), "households")
```
Render to HTML, Word and PDF
quarto preview profile.qmd # render and open in a browser, re-rendering on save quarto render profile.qmd --to html # a single HTML file quarto render profile.qmd --to docx # a Word document quarto render profile.qmd --to pdf # a PDF (needs a TeX installation, below) quarto install tinytex # the TeX distribution Quarto recommends for PDF
The Quarto tutorial notes that the file name should come first after quarto render. For PDF, the PDF guide says you need a TeX distribution and recommends TinyTeX, installed with the last command above. Word output needs nothing extra.
profile.qmd in a folder with households.csv, run quarto render profile.qmd --to html, and open the HTML file. Look for the sentence "Households surveyed: 240." and a table of ten district percentages. Then render --to docx and open the Word file: the same numbers, no copying.Parameterised reports: one template, ten district profiles
A parameter is a value the report reads from outside, such as the district name. Write the report once, then render it once per district. The Quarto parameters guide gives one syntax for each engine.
- R (knitr): declare
params:in the YAML header and useparams$districtin code. - Python (Jupyter): tag one cell
parametersand set default values there; Quarto injects a cell after it with the values you pass.
---
title: "District profile: `r params$district`"
format: html
params:
district: "Purnia"
---
```{r}
hh <- read.csv("households.csv")
g <- hh[hh$district == params$district, ]
```
`r params$district` has `r nrow(g)` surveyed households.
```{python}
#| tags: [parameters]
district = "Purnia"
```
# One district quarto render profile.qmd -P district:Gaya --output gaya.html # Every district, in a terminal (macOS, Linux, or Git Bash on Windows) for d in Barmer Betul Gaya Indore Kozhikode Patna Purnia Rewa Udaipur Wayanad; do quarto render profile.qmd -P district:$d --output profile-$d.html done
-P district:Gaya sets the parameter (the guide's own example is -P alpha:0.2) and --output names the file, so the ten renders do not overwrite each other.
The cell runs what the report's chunk computes, for one district. Switch between R and Python with the tabs on the cell.
For Purnia (Bihar) it reports 24 households, 75.0% with a toilet, 75.0% with a bank account and mean monthly spending of Rs 2,395 per person, then the count of households in each caste group. Your rendered Purnia profile should show the same figures.
Kozhikode and run again; then to Kozhikkode. The misspelt name finds no households: R prints 0 households and NaN percentages, and Python stops with an error when it looks up the state. Add a check at the top of your .qmd that stops the render when the district is not in districts.csv, so a typo in a loop cannot produce an empty profile with a confident title.A reproducible project layout
A project is reproducible when someone else can take the repository and the raw data, run it, and get the same report. A plain folder layout and three rules get most of the way.
district-profiles/
README.md what this is, how to run it, who to ask
.gitignore data/raw/, outputs/, credentials
data/
raw/ exports exactly as downloaded: never edited, never committed
clean/ written only by the cleaning script
scripts/
01_clean.R raw -> clean, with checks that stop on a problem
analysis/
profile.qmd the parameterised report
outputs/ rendered reports: rebuilt, never edited by hand
- Raw data is read-only. Scripts read
data/raw/and writedata/clean/; nobody edits a raw export by hand. Raw data stays out of Git and is shared through your organisation's controlled storage. - Outputs are rebuilt. Anything in
outputs/can be deleted and regenerated with the render loop, so it does not need to be committed. - Numbered scripts run in order.
01_clean.Rbefore the report. The README says which command to run. - Record versions. Note the R or Python version and package versions in the README (R's
sessionInfo()prints them), so a result can be traced to the software that produced it. - Commit small and often, each commit doing one thing, with a message that says what.
Where next
R and Python side by side
The language for the code chunks in your reports.
The tidyverse
Write the tables and charts your Quarto profiles will hold.
Spreadsheets for M&E
Where most of the numbers start, and when to leave them.
Data Protection and the DPDP Act
What the law asks before personal data goes anywhere.
Research Ethics 101
Consent, confidentiality and data handling in field research.