Open Data Editor: checking a dataset before you share it
Before a survey file or MIS extract goes to a partner, a portal or a funder, check that its structure holds: one name per column, the same number of cells in every row, numbers where numbers belong, and the rules you promised (unique IDs, allowed values). Open Data Editor does these checks without code, on your own computer. Python cells on this page run the same checks so you can see what each error means.
What Open Data Editor is, and its status
Open Data Editor (ODE) is made by the Open Knowledge Foundation (OKFN). Its documentation describes it as "a free, open-source tool designed to help nonprofits, data journalists, activists, and public servants detect errors in their datasets", for people who work with tables in Excel, Google Sheets or CSV and do not write code. It checks files against the rules of the Frictionless framework, an open standard for describing tables, and the OKFN page says the Digital Public Goods Alliance has recognised it as a digital public good since 2025.
What it does well
- Finds structural errors as soon as you open a file: missing or duplicate column names, empty rows, rows with too few or too many cells, and cells of the wrong type.
- Checks rules you add as metadata: a column must be filled in, unique, one of a list of values, within a range, or match a pattern.
- Lets you fix cells in a grid and re-check, then export the file, or an Excel workbook listing every error.
- Works locally. Its documentation calls it local-first, and the optional AI assistant uses a model downloaded to your computer; the documentation says no data is sent to the cloud.
What it is not
It is not a cleaning tool in the OpenRefine sense: it shows you that Purnea and Purnia both appear only if you have told it which values are allowed, and it does not cluster or bulk-merge them. Use OpenRefine to clean, and Open Data Editor to check the result before it leaves your hands.
Download and install
The download guide offers one file per system, from the OKFN project page or from GitHub releases:
- Windows: the most recent EXE file.
- macOS: the most recent DMG file.
- Linux: an AppImage (any distribution) or a DEB file (Ubuntu and Debian).
- Windows: download the EXE. If the browser warns about the download, choose to keep it (the guide shows a "Continue download" prompt). Double-click it. If Windows shows a security window, click More info, then Run anyway.
- macOS: open the DMG. If macOS shows a security message, the guide says to click the question mark in it, follow the link in the first section, and change the setting that allows the app to run.
- Linux AppImage: make the file executable (in a terminal,
chmod +xfollowed by the file name), then double-click it. - Ubuntu or Debian DEB: double-click it, or install it from a terminal with the command below.
# Replace <version> with the version you downloaded sudo dpkg -i opendataeditor-linux-<version>.deb
Open a CSV
Download the two course files: households.csv (240 households) and districts.csv (10 districts). Illustrative data, invented for teaching The district names are real places; every number is made up.
- Open Open Data Editor and click Upload your data (in the centre of the start screen, or at the top left of the sidebar).
- Under From your computer, choose Add one or more Excel or csv files and select
households.csv. To check a whole folder of files, use Add one or more folders. - The file appears in the sidebar. Click it: ODE validates it and shows it in the grid.
Tables published online
Under Add External Data you can paste the address of a table on an open data portal, a Google Sheet or a GitHub repository. For a Google Sheet, the guide says the sheet must be published to the web and you should paste its normal address, the one ending in /edit. The guide marks the published pubhtml address as the wrong one to use.
Prepare the file before you open it
- One header row of column names, starting in the first row. Remove titles, notes, logos and totals above or beside the table; the guide warns that these produce many errors.
- No merged cells.
- Decimals written with a full stop. The guide warns that a comma as decimal separator makes the number column read as text.
The errors it reports
A file with problems gets a red dot beside its name in the sidebar, and problem cells are shaded red in the grid. Click Errors Report at the top left of the grid for the full list. The full list of errors names seven it finds without any metadata:
- Header missing (Blank Label): the header row is empty.
- Column name missing: one or more columns have no name.
- Duplicate column name: two columns share a name.
- Empty row: a row with no values.
- Missing cell: a row has fewer cells than the header.
- Extra cell: a row has more cells than the header.
- Wrong data type: a cell does not fit the column's type, such as text in a column of numbers. The guide notes ODE can only infer the intended type when enough cells in the column already have it.
The cell below builds a small broken file with five of those problems, checks it with a short function written for this course (not ODE's own code), and then checks the real households.csv.
The broken file produces seven errors: column 4 has no name, area appears twice (columns 3 and 6), row 4 is empty, row 5 is missing a cell, row 6 has an extra one, and in row 7 "2,080" is text in a column of whole numbers. households.csv has 0 errors. To see the same file in ODE, copy the eight lines between the triple quotes into a text editor, save them as broken.csv and upload it: its Errors Report should point to the same rows and columns.
,,,,, and run again. Which error disappears, and which row numbers move?Fix errors and re-check
For a handful of bad cells, fix them in ODE's grid, following the editing guide:
- Find the red cell (the Errors Report lists its row and column).
- Double-click it and type the corrected value.
- Click elsewhere in the table to accept the change.
- Click Save changes, which becomes active when there are unsaved edits. ODE re-validates and updates the Errors Report.
Know when to stop editing by hand
- Structural errors (missing or extra cells, a blank or duplicate column name) usually come from how the file was exported: a stray comma inside an unquoted text field, a notes column, two tables in one sheet. Fix the export and export again; patching rows by hand leaves the cause in place for next month.
- Many wrong-type cells in one column ("2,080", "NA", "not asked") are a cleaning job. Do it in OpenRefine or in code, where the step is recorded and can be repeated.
- A wrong value that is the right type, such as a household size of 0, passes every check in module 4. Only a rule catches it, which is what module 6 adds.
Metadata and rules
Metadata is the description that travels with the file: a title, what each column means, its type, and the rules its values must follow. In ODE, select a file and click any cell of the header row to open the Metadata window; edit, then click Save changes, which re-validates the file.
The rules follow the Table Schema standard. ODE's error guide says it supports these field constraints: required, enum (a list of allowed values), minimum, maximum, minLength, maxLength and pattern, plus unique, a primary key and foreign keys. With them it can report seven more errors: extra, missing and incorrect column names against the schema, and primary key, foreign key, unique and constraint errors.
The cell below writes rules of that kind for the household file and applies them, first to households.csv and then to a copy with five mistakes typed in.
The real file breaks no rule. The copy breaks all five, and each one is a mistake that no structural check sees: a repeated ID, Gaya with a trailing space (not in districts.csv), rural in lower case, a household of 0 people, and a blank expenditure.
The rules as a Table Schema
Written in the standard's JSON form, the same rules look like this. A partner with any Frictionless tool can check your file against it.
> 20 to > 6 and run. The real file now fails on every household of 7 or 8 people. Those households are real; the rule was wrong. Set limits from what is possible in the population; the largest value in one file is a poor guide.Export, share and publish
Click Export at the top right of the grid. The export guide lists two options:
- Download file: the table as CSV.
- Download file with errors: an Excel workbook with three sheets. Data is the table with errors painted red, Errors Description lists each error with its row and column, and Blank Rows holds the rows that had no values. Send this to the person who produced the file: it tells them exactly what to fix.
Publishing
The README describes ODE as an application to "explore, validate and publish data", but the current user guide (checked 6 October 2026) has no section on publishing to a data portal: its last step is export. So publish the exported CSV through your portal's own upload page (a CKAN portal, Zenodo, your organisation's repository), and attach the data dictionary and the schema. If your installed version shows a publish option, check its destination before you use it.
The dictionary has one row per column of the household file (13 rows), with each column's type, its count of missing values (0 in this file), how many distinct values it holds and an example. Save it next to the CSV.
A checklist before the file leaves your computer
- The Errors Report is empty, or every remaining error is explained in a note.
- No column holds names, phone numbers, Aadhaar numbers, exact addresses or GPS points unless sharing them is allowed. India's Digital Personal Data Protection Act 2023 governs personal data in digital form; check your consent notice and data-sharing agreement before you share.
- Small cells are safe: a table of caste by village with 1 or 2 households in a cell can identify families even without names.
- The data dictionary and the schema go with the file, along with who collected it, when, and the licence you are publishing under.
Where next
OpenRefine: cleaning messy survey and MIS data
Clean the file first: facets, clustering and a replayable recipe.
QGIS: maps for development data
Put the checked district file on a map.
pandas for development data
Write checks like these as code you can run on every new export.
Data Protection and the DPDP Act
What you may share, with whom, and on what basis.
Data Literacy 101
Why a well-described table is worth more than a large one.