Initiatives
Documentary films
We have made 10 documentary films since 2018. Elder Gyani Maiya Sen Kusunda narrates Gyani Maiya (2019) in the Kusunda language of Nepal. Mage Porob follows a Ho village holding its harvest festival amid mining and land loss, and Remosam the Bondak people of the Bonda hills in Odisha. In Nani Ma (2022), 95-year-old Musamoni Panigrahi shares songs and stories shaped by the 1866 Orissa famine. The Volunteer Archivists (2022) follows the people digitising two centuries of printed Odia books.
OpenSpeaks
OpenSpeaks, a community collective to document and archive languages. It builds capacity, co-documents audio-visual media, and makes free and open-source software, frameworks and workflows with speakers of endangered, Indigenous and other local languages. Between 2017 and 2026, it has documented over 20 South Asian languages.
Recordings are archived at the OpenSpeaks Archives, the Endangered Languages Archive, the Language Archive Cologne (bundle 1, bundle 2) and the Library of Congress.
Archiving the Present
Archiving the Present is a community network for archiving oral languages, memories and knowledges, organised by OpenSpeaks. Between 18 May and 24 August 2026, it ran eight online sessions for 101 applicants from 14 countries, working in about 115 languages.
The network launched on 3 September 2026 at TinkerSpace in Kochi, with 34 people. Nineteen of them are community documenters and archivists working in 27 languages and varieties. The sessions became a manual, Archiving the Present: A Community Guide to Archiving Languages, Memories, and Knowledges, in print and as an e-book, with cover art by Siddhesh Gautam, and a free course on WikiLearn. Samagata Foundation, The Wikimedia Foundation, Creative Commons supported the network.
Speech data
OpenSpeaks Voice: Odia holds over 25 hours of Odia speech under CC0. Nearly 66,000 words (about 22 hours) are on Wikimedia Commons, recorded with Lingua Libre, and 8 hours of sentences are on Mozilla Common Voice. About 600 recordings are by the late Musamoni Panigrahi in Baleswari Odia, taken from the archival footage of Nani Ma. Subhashish Panigrahi recorded the rest in Baleswari and Mugalbandi Odia. The Library of Congress catalogued the dataset in 2024 as OpenSpeaks Data Pages, then the largest public-domain speech dataset in Odia. The method is in the paper Building a Public Domain Voice Database for Odia (The Web Conference 2022).
Registry of Type Design
The Registry of Type Design records typefaces and the people who made them, digital and pre-digital, in any script. It holds 167 typefaces and 143 people, including Odia, Ol Chiki and Warang Citi type. Every record has a permanent id and is open under CC BY-SA 4.0 as web pages, a JSON API and CSV downloads. Wikidata links to it through property P14791. The data is on GitHub.
Type design
Chapakala 19 revives the Odia letterpress type of an 1875 Oriya New Testament. Subhashish Panigrahi designed it between 2024 and 2026, with advice from Nasim Ali, Yesha Goshar and Liang Hai. The font is under the SIL Open Font License. Its letterforms trained the historical Odia OCR model below.
Tools
Subtitler
Subtitler subtitles videos from Wikimedia Commons or your own files. It finds the pauses in speech and splits the recording into draft subtitles to edit.
Tome
Tome writes metadata for oral knowledge recordings and generates consent agreements.
Bento
Bento is a media toolkit for oral knowledge archives. It runs in the browser and works offline.
Text recognition (OCR)
We train Tesseract models so that printed text in our languages can be searched and reused. They are open in ofdn/tessdata_contrib.
- Santali (Ol Chiki): the first LSTM model for Ol Chiki, trained with the Noto Sans Ol Chiki and Guru Gomke fonts.
- Historical Odia: for Odia printed in nineteenth- and early twentieth-century letterpress. The standard Odia model, last updated in 2017, fails on this type. Ours is fine-tuned on 5,800 lines set in Chapakala 19.
- Ho (Warang Citi): in training. Warang Citi is the script Lako Bodra created for Ho.
Santali Unicode Encoding Converter
The Santali Unicode Encoding Converter turns Santali text typed in ASCII and other legacy encodings into Unicode Ol Chiki. Jnanaranjan Sahu wrote it; it is under the MIT License.
Earlier work
- OFDN Conversations: a podcast with activists, artists and technologists. Six episodes, 2020 to 2023.
- Language resources: media and information literacy resources in Santali, Odia, Ho and Mihaq (Kusunda), 2018 to 2021.
- Pothi: documentation of intangible cultural heritage, led by Prateek Pattanaik, until 2021.
- Marginalized Community Council: an online working group of activists, started in 2019 and coordinated by Ramjit Tudu.
- Events: conferences and workshops we organised or took part in, 2017 to 2018.