Skip to content
ImpactMojo ImpactMojo
Browse Membership
Back to Code Studio
Code Studio · Guided tool course

Open Data Editor: checking a dataset before you share it

Before a survey file or MIS extract goes to a partner, a portal or a funder, check that its structure holds: one name per column, the same number of cells in every row, numbers where numbers belong, and the rules you promised (unique IDs, allowed values). Open Data Editor does these checks without code, on your own computer. Python cells on this page run the same checks so you can see what each error means.

Open Data Editor is a desktop application. It does not run on this page. The steps tell you what to click and what to look for; this page does not show Open Data Editor screens, because it cannot produce them. The Python cells run equivalent checks here with pandas and Python's csv module, so you can see each kind of error on a small file. The first Python run downloads the engine once (about 10 MB).
Module 1 of 7

What Open Data Editor is, and its status

Open Data Editor (ODE) is made by the Open Knowledge Foundation (OKFN). Its documentation describes it as "a free, open-source tool designed to help nonprofits, data journalists, activists, and public servants detect errors in their datasets", for people who work with tables in Excel, Google Sheets or CSV and do not write code. It checks files against the rules of the Frictionless framework, an open standard for describing tables, and the OKFN page says the Digital Public Goods Alliance has recognised it as a digital public good since 2025.

Version and status, checked 6 October 2026. The latest release on ODE's GitHub releases page is v1.8.0, published 17 September 2026; the previous one, v1.7.1, came out on 23 October 2025. The user guide at opendataeditor.okfn.org is headed "Open Data Editor 1.5.1 documentation", so it trails the software by several releases: if a button on your screen differs from these steps, trust your screen. The repository README still titles the project "Open Data Editor (beta)". The source code is under the MIT licence and the application is free.

What it does well

What it is not

It is not a cleaning tool in the OpenRefine sense: it shows you that Purnea and Purnia both appear only if you have told it which values are allowed, and it does not cluster or bulk-merge them. Use OpenRefine to clean, and Open Data Editor to check the result before it leaves your hands.

Module 2 of 7

Download and install

The download guide offers one file per system, from the OKFN project page or from GitHub releases:

  1. Windows: download the EXE. If the browser warns about the download, choose to keep it (the guide shows a "Continue download" prompt). Double-click it. If Windows shows a security window, click More info, then Run anyway.
  2. macOS: open the DMG. If macOS shows a security message, the guide says to click the question mark in it, follow the link in the first section, and change the setting that allows the app to run.
  3. Linux AppImage: make the file executable (in a terminal, chmod +x followed by the file name), then double-click it.
  4. Ubuntu or Debian DEB: double-click it, or install it from a terminal with the command below.
Terminal (Debian or Ubuntu), from the ODE install guide
# Replace <version> with the version you downloaded
sudo dpkg -i opendataeditor-linux-<version>.deb
On a modest laptop. The checks themselves are light. The optional AI assistant is not: the GitHub README says its default model, Apertus 8B (Apache 2.0 licence), "requires at least 6GB of video memory", and the model is a separate download that ODE offers the first time you press the AI button. Skip it on a machine without a graphics card; nothing else depends on it.
Module 3 of 7

Open a CSV

Download the two course files: households.csv (240 households) and districts.csv (10 districts). Illustrative data, invented for teaching The district names are real places; every number is made up.

  1. Open Open Data Editor and click Upload your data (in the centre of the start screen, or at the top left of the sidebar).
  2. Under From your computer, choose Add one or more Excel or csv files and select households.csv. To check a whole folder of files, use Add one or more folders.
  3. The file appears in the sidebar. Click it: ODE validates it and shows it in the grid.
ODE works on a copy. The upload guide says ODE does not change your original file: it copies it into the application's own folder. Right-click the file in the sidebar and choose Open Location to see where the copy lives. Your edits go to that copy, so export it when you are done (module 7).

Tables published online

Under Add External Data you can paste the address of a table on an open data portal, a Google Sheet or a GitHub repository. For a Google Sheet, the guide says the sheet must be published to the web and you should paste its normal address, the one ending in /edit. The guide marks the published pubhtml address as the wrong one to use.

Prepare the file before you open it

Module 4 of 7

The errors it reports

A file with problems gets a red dot beside its name in the sidebar, and problem cells are shaded red in the grid. Click Errors Report at the top left of the grid for the full list. The full list of errors names seven it finds without any metadata:

The cell below builds a small broken file with five of those problems, checks it with a short function written for this course (not ODE's own code), and then checks the real households.csv.

The broken file produces seven errors: column 4 has no name, area appears twice (columns 3 and 6), row 4 is empty, row 5 is missing a cell, row 6 has an extra one, and in row 7 "2,080" is text in a column of whole numbers. households.csv has 0 errors. To see the same file in ODE, copy the eight lines between the triple quotes into a text editor, save them as broken.csv and upload it: its Errors Report should point to the same rows and columns.

Exercise. In the cell, delete the line ,,,,, and run again. Which error disappears, and which row numbers move?
Module 5 of 7

Fix errors and re-check

For a handful of bad cells, fix them in ODE's grid, following the editing guide:

  1. Find the red cell (the Errors Report lists its row and column).
  2. Double-click it and type the corrected value.
  3. Click elsewhere in the table to accept the change.
  4. Click Save changes, which becomes active when there are unsaved edits. ODE re-validates and updates the Errors Report.

Know when to stop editing by hand

Do not invent a value to clear an error. If an expenditure cell says "not asked", the honest fix is an empty cell (missing). Do not type a zero or the column mean. A clean report on a file with made-up values is worse than a report with errors.
Module 6 of 7

Metadata and rules

Metadata is the description that travels with the file: a title, what each column means, its type, and the rules its values must follow. In ODE, select a file and click any cell of the header row to open the Metadata window; edit, then click Save changes, which re-validates the file.

The rules follow the Table Schema standard. ODE's error guide says it supports these field constraints: required, enum (a list of allowed values), minimum, maximum, minLength, maxLength and pattern, plus unique, a primary key and foreign keys. With them it can report seven more errors: extra, missing and incorrect column names against the schema, and primary key, foreign key, unique and constraint errors.

The cell below writes rules of that kind for the household file and applies them, first to households.csv and then to a copy with five mistakes typed in.

The real file breaks no rule. The copy breaks all five, and each one is a mistake that no structural check sees: a repeated ID, Gaya with a trailing space (not in districts.csv), rural in lower case, a household of 0 people, and a blank expenditure.

The rules as a Table Schema

Written in the standard's JSON form, the same rules look like this. A partner with any Frictionless tool can check your file against it.

Exercise. In the rules cell, change > 20 to > 6 and run. The real file now fails on every household of 7 or 8 people. Those households are real; the rule was wrong. Set limits from what is possible in the population; the largest value in one file is a poor guide.
Module 7 of 7

Export, share and publish

Click Export at the top right of the grid. The export guide lists two options:

Publishing

The README describes ODE as an application to "explore, validate and publish data", but the current user guide (checked 6 October 2026) has no section on publishing to a data portal: its last step is export. So publish the exported CSV through your portal's own upload page (a CKAN portal, Zenodo, your organisation's repository), and attach the data dictionary and the schema. If your installed version shows a publish option, check its destination before you use it.

The dictionary has one row per column of the household file (13 rows), with each column's type, its count of missing values (0 in this file), how many distinct values it holds and an example. Save it next to the CSV.

A checklist before the file leaves your computer

Where next

→

OpenRefine: cleaning messy survey and MIS data

Clean the file first: facets, clustering and a replayable recipe.

→

QGIS: maps for development data

Put the checked district file on a map.

→

pandas for development data

Write checks like these as code you can run on every new export.

→

Data Protection and the DPDP Act

What you may share, with whom, and on what basis.

→

Data Literacy 101

Why a well-described table is worth more than a large one.