Third-party notices
Credit for the third-party data that went into Falcon models: what each set is, who made it, its licence and a link to it.
Four to six sentences from Cognitio, written on request. Generated text: the page is the reference.
On this page
Most Falcon models are trained only on data that Ducky Software generates itself. Where a model was trained with data made by others, this page gives the credit that data's licence asks for: the name of the set, its authors, its licence and where to find it. The same credit is written into the model's own metadata (meta.json, field third_party_data), which ships with the weights.
Using a set here does not mean its authors endorse Falcon or any Falcon model. None of the sets below was changed except as described: rows were selected and filtered, never rewritten. The full text of each licence named here, and of the licences of the other models, fonts and software Falcon uses, is on Licenses.
ELIM and LIM Nano (1.1.0) #
From version 1.1.0 (2026-10-03), ELIM and LIM Nano are trained on synthetic data and licence-clean public utterances. The public utterances are ordinary, harmless requests (weather, timers, music, travel and the like). They are used only as examples of messages that should be allowed: no unsafe example comes from them. Utterances that match a list of risk words were dropped before training, and so was the fraud-reporting intent of CLINC150. Only text was used. LIM and the 1.0.0 versions of ELIM and LIM Nano were not trained on any of these sets.
| Set | Authors | Licence | Source | What was used |
|---|---|---|---|---|
| CLINC150, from “An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction” | Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang and Jason Mars (EMNLP-IJCNLP 2019) | CC BY 3.0 | github.com/clinc/oos-eval, revision 828f809 | Queries from the train, validation and out-of-scope train and validation splits |
| MASSIVE 1.0, English (US), from “MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages” | Jack FitzGerald and co-authors, Amazon (2022) | CC BY 4.0 | github.com/alexa/massive | English (US) utterances from the train and development splits |
| SLURP, from “SLURP: A Spoken Language Understanding Resource Package” (the English text that MASSIVE builds on) | Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski and Verena Rieser (EMNLP 2020) | CC BY 4.0 for the text | github.com/pswietojanski/slurp | The text of the MASSIVE English (US) utterances above; no audio |
Voice data #
The Express voice models were trained with recorded speech published under open licences. Their credit is given on their own pages, where it has been since they were released: Express Voice and Express Voice Blend.
Questions #
Ducky Software keeps a record of each set's source, revision and checksums. If you think a credit here is missing or wrong, please tell Ducky Software; the Access page says how access to the preview is arranged.