msmth commited on
Commit
38c2e17
·
verified ·
1 Parent(s): c904ce8

Pre-release review fixes: code robustness, card corrections

Browse files
Files changed (6) hide show
  1. CREDITS_BOOKS.tsv +27 -25
  2. EVALUATION.md +59 -29
  3. NOTICE +21 -14
  4. README.md +50 -18
  5. requirements.txt +7 -4
  6. source1.py +372 -88
CREDITS_BOOKS.tsv CHANGED
@@ -1,8 +1,10 @@
1
- # Source-1: per-work credits for the books in its training and validation splits that are not public domain or
2
- # CC0 (2,846 works). Each work is credited to the authors and publishers named in it, under the licence recorded
3
- # for it by the platform it came from (for three works, the IGO licence their own text states; the licence cell
4
- # says so). Modified: used to train a classifier. No text of these works is distributed with Source-1. This file
5
- # is part of NOTICE (see NOTICE, part 3).
 
 
6
  # citation_or_note: the citation or attribution that the work's own text asks for, as the work gives it (line
7
  # breaks joined, PDF spacing repaired), for FAO books, OpenStax textbooks, Eurydice, JRC and other EU reports, and
8
  # some other books and reports (university-press books, Frontiers ebooks, research and project reports); or a
@@ -107,7 +109,7 @@ DOAB Prácticas lingüísticas heterogéneas: Nuevas perspectivas para el estudi
107
  DOAB Questioni di donne: Diplomazia informale e reti femminili alla corte dei Savoia-Carignano (XVII secolo) Lurgo, Elisabetta it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://directory.doabooks.org/handle/20.500.12854/134564
108
  DOAB Religion og etikk i skole og barnehage Afset, Bente; Redse, Arne no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://directory.doabooks.org/handle/20.500.12854/38379
109
  DOAB Réinventer l’art sacré: Le Groupe de Saint-Luc (1919-1945) Noverraz, Camille fr CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://directory.doabooks.org/handle/20.500.12854/152369
110
- DOAB Sjezd českých právníků 2022 cs CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://directory.doabooks.org/handle/20.500.12854/98695
111
  DOAB Slovanský literární svět: kontexty a konfrontace III: Motiv domova ve slovanských literaturách Bujnáková, Jana; Cepková Feješová, Zuzana; Derková, Vladimíra; Eniko, Mateja; Heinigová, Lenka; Hrancová, Hana cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80840
112
  DOAB Somatopedické simulační techniky a intervence: Metodické texty k projektu MUNI 4.0. Pedagogická fakulta, studijní program Logopedie (Bc.) Opatřilová, Dagmar cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80845
113
  DOAB Speciálněpedagogická diagnostika somatopedická: Metodické texty k projektu MUNI 4.0. Pedagogická fakulta, studijní program Logopedie (Bc.) Opatřilová, Dagmar cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80846
@@ -767,7 +769,7 @@ FAO Combattre la criminalité liée aux forêts en Afrique de l'Ouest Chasi, R.M
767
  FAO Commission générale des pêches pour la Méditerranée: Rapport de la quarante-huitième session FAO; fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd9408fr Citation requested in the work: FAO. 2026. Commission générale des pêches pour la Méditerranée: Rapport de la quarante-huitième session – Malaga, Espagne, 4-9 novembre 2025. Commission générale des pêches pour la Méditerranée (CGPM) – Rapports de session, n°48. Rome. https://doi.org/10.4060/cd9408fr
768
  FAO Compendium of decisions made by the Parties to the FAO Agreement on Port State Measures FAO; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0286en Citation requested in the work: FAO. 2026. Compendium of decisions made by the Parties to the FAO Agreement on Port State Measures. Second edition. Rome. https://doi.org/10.4060/ce0286en
769
  FAO Comptes rendus du deuxième Sommet sur la pêche artisanale FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3997fr Citation requested in the work: FAO. 2025. Comptes rendus du deuxième Sommet sur la pêche artisanale, 5-7 juillet 2024, Rome. FAO Comptes rendus des pêches et de l’aquaculture, n° 70. Rome. https://doi.org/10.4060/cd3997fr
770
- FAO Cостояние мирового рыболовства и аквакультуры – 2026 ​ФАО; ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd8357ru Citation requested in the work: "ФАО. 2026. Cостояние мирового рыболовства и аквакультуры – 2026. ""Голубая трансформация"": от замысла к практическим результатам. Рим. https://doi.org/10.4060/cd8357ru"
771
  FAO Design of a climate-proof fish buying station Josupeit, H.; Moretti, S.; Sciortino, J.A.; Urbani, R.; van Anrooy, R.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1081en Citation requested in the work: Josupeit, H., Moretti, S., Sciortino, J.A., Urbani, R. & Van Anrooy, R. 2026. Design of a climate-proof fish buying station. FAO Fisheries and Aquaculture Technical Paper, No. 691. Rome, FAO. https://doi.org/10.4060/ce1081en
772
  FAO Developing holistic nutrition guidelines and standards for school meals FAO; WFP; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1121en Citation requested in the work: FAO and WFP. 2026. Developing holistic nutrition guidelines and standards for school meals – A global methodology. Rome. https://doi.org/10.4060/ce1121en
773
  FAO Developing nutrition-sensitive value chains Andrianarimanana, M.; Galante, A.; Liu, B.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0611en Citation requested in the work: Andrianarimanana, M., Galante, A. & Liu, B. 2026. Developing nutrition‑sensitive value chains – Guidelines for practitioners. Rome, FAO. https://doi.org/10.4060/ce0611en
@@ -818,7 +820,7 @@ FAO Roles and values of camelids and their products FAO; en CC BY 4.0 https://op
818
  FAO Scaling up community-base fisheries management in the Pacific Govan, H.; Tuxson, T.; Tauati, M.; Lalavanua, W.; Schwarz, A.; Kinch, J.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1087en Citation requested in the work: Govan, H., Tuxson,T., Tauati, M., Lalavanua, W., Schwarz, A. & Kinch, J. 2026. Scaling up community-based fisheries management in the Pacific – Outlook and prospects for securing sustainable coastal fisheries, livelihoods and ecosystems. Rome, FAO; Noumea, Pacific Community. https://doi.org/10.4060/ce1087en
819
  FAO Seed to sip: Arabica coffee production guide for Saudi Arabia Gichimu, B.M.; Ghosh, K.; Alfaifi, B.H.; Alfaifi, K.A.; Al Mutlaq, A.M.; Bustamante Adum, D.; Lubabali, H.A.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0725en Citation requested in the work: Gichimu, B.M, Ghosh, K., Alfaifi, B.H., Alfaifi, K.A., Al Mutlaq, A.M., Bustamante Adum, D. & Lubabali, H.A. 2026. Seed to sip: Arabica coffee production guide for Saudi Arabia. Riyadh, FAO. https://doi.org/10.4060/ce0725en
820
  FAO Shaping agrifood systems legislation Rosenbaum, K.L.; Vidar, M.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1406en Citation requested in the work: Rosenbaum, K.L and Vidar, M. 2026. Shaping agrifood systems legislation – Good practices for legal advisors in drafting laws and supporting related processes. Legal Guide No. 5. Rome, FAO. https://doi.org/10.4060/ce1406en
821
- FAO Soil health and fertilizer Jafari, A.; Le Cotty, T.; Tefft, J.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd9916en Citation requested in the work: Jafari, A., Le Cotty, T. & Tefft, J. 2026. Soil health and fertilizer – Policy and investment prospects in sub-Saharan Africa. Directions in Investment no. 19. Rome, Paris and Brussels, FAO, Agrinatura and the European Union. https://doi. org/10.4060/cd9916en
822
  FAO Status of the world's soil resources 2026 FAO; ITPS; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0792en Citation requested in the work: FAO and ITPS. 2026. Status of the world's soil resources 2026. Rome, FAO. https://doi.org/10.4060/ce0792en
823
  FAO Strengthening small and medium agroenterprise finance in the Near East and North Africa region Aldredge, H.; Priebe, J.; Zook, D.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0863en Citation requested in the work: Aldredge, H., Priebe, J. & Zook, D. 2026. Strengthening small and medium agroenterprise finance in the Near East and North Africa region: Bridging the finance gap. Rome, FAO. https://doi.org/10.4060/ce0863en
824
  FAO Sổ tay hỏi đáp Phùng, T.V.; Tôn, V.Đ.; Sơn, T.H; Phục, N.N.; vi CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0930vi Citation requested in the work: Phùng, T.V., Tôn, V.Đ, Sơn, T.H. and Phục, N.N. 2026. Sổ tay hỏi đáp về th c h nh t t an to n sinh học v xử lý chất thải trong chăn nuôi lợn quy mô vừa v nhỏ. Hà Nội, FAO.
@@ -832,14 +834,14 @@ FAO Vodič za odgovornu i racionalnu primenu antibiotika kod goveda FAO; sr CC B
832
  FAO Vodič za odgovornu i racionalnu primenu antibiotika kod ovaca i koza Krnjaić, D.; Savić, B.; Trailović, S.; sr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0604sr Citation requested in the work: Krnjaić, D, Savić, B, Trailović, S. 2026. Vodič za odgovornu i racionalnu primenu antibiotika kod ovaca i koza. Beograd, FAO.
833
  FAO Vodič za odgovornu i racionalnu primenu antibiotika kod pasa i mačaka Krnjaić, D.; Resanović, R,; sr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0991sr Citation requested in the work: Krnjaić, D, Resanović, R, Trailović, S. 2026. Vodič za odgovornu i racionalnu primenu antibiotika kod pasa i mačaka . Beograd, FAO.
834
  FAO Voices of cooperators: A collective journey through the agrifood systems in the Near East and North Africa FAO; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0708en Citation requested in the work: FAO. 2026. Voices of cooperators: A collective journey through the agrifood systems in the Near East and North Africa region. Cairo. https://doi.org/10.4060/ce0708en
835
- FAO Wood products in the bioeconomy Reck, B.K.; Johnston, C.; Foong, A.; Gupta, A.; Karpov, A.; Holsten, A.; Misselwitz, P.; Keenan, R.J.; Formenton Cardoso, N.; Walter, S.; Bull, L.; Steel, E.A.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0315en Citation requested in the work: Reck, B.K., Johnston, C., Foong, A., Gupta, A., Karpov, A., Holsten, A., Misselwitz, P., Keenan, R.J., Formenton Cardoso, N., Walter, S., Bull, L., and Steel, E.A. 2026. Wood products in the bioeconomy: Scenario-based assessment of the potential for engineered wood products in climate change mitigation. Rome, FAO. DOI https:// doi.org/10.4060/ce0315en
836
  FAO Évaluation du projet «Consolidation de la paix dans la zone frontalière du nord-est de la Côte d’Ivoire par le biais d’une gestion participative et planifiée des ressources naturelles» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3911fr Citation requested in the work: FAO. 2025. Évaluation du projet «Consolidation de la paix dans la zone frontalière du nord-est de la Côte d’Ivoire, par le biais d’une gestion participative et planifiée des ressources naturelles» – Code du projet: UNJP/IVC/037/PBF. Série évaluation de projet, 01/2025. Rome. https://doi.org/10.4060/cd3911fr
837
  FAO Évaluation du projet «Promouvoir une production de cacao sans déforestation pour réduire les émissions en Côte d’Ivoire» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3686fr Citation requested in the work: FAO. 2025. Évaluation du projet «Promouvoir une production de cacao sans déforestation pour réduire les émissions en Côte d’Ivoire» Rapport de mi-parcous, code du projet: GCP/IVC/609/GCF. Série évaluation de projet, n° 50/2024. Rome. https://doi.org/10.4060/cd3686fr
838
  FAO Évaluation du projet «Soutien à l’auto-emploi de la jeunesse rurale, vecteur de paix et de cohésion sociale au Mali» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7607fr Citation requested in the work: FAO. 2025. Évaluation du projet «Soutien à l’auto-emploi de la jeunesse rurale, vecteur de paix et de cohésion sociale au Mali» Code du projet: UNJP/MLI/068/PBF. Série évaluation de projet, n. 26/2025. Rome. https://doi.org/10.4060/cd7607fr
839
  FAO Évaluation finale du projet «Revitaliser les agroécosystèmes oasiens à travers une approche durable, intégrée et paysagère dans la région de Drâa Tafilalet» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5845fr Citation requested in the work: FAO. 2025. Évaluation finale du projet «Revitaliser les agroécosystèmes oasiens à travers une approche durable, intégrée et paysagère dans la région de Drâa Tafilalet» - Code du projet: GCP/MOR/046/GFF. Série évaluation de projet, N.° 15/2025. Rome. https://doi.org/10.4060/cd5845fr
840
  FAO Évaluation finale du projet «Élimination des polluants organiques persistants et pesticides obsolètes et renforcement de la gestion rationnelle des pesticides au Cameroun» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5148fr Citation requested in the work: FAO. 2025. Évaluation finale du projet «Élimination des polluants organiques persistants et pesticides obsolètes et renforcement de la gestion rationnelle des pesticides au Cameroun» – Code du projet: GCP/CMR/031/GFF, Identifiant FEM: 4641. Série Évaluation de projet, 12/2025 Rome. https://doi.org/10.4060/cd5148fr
841
- FAO Анализ кооперативного законодательства Российской Федерации и других стран СНГ Клименко, О.И.; Кондракова, И.А.; Мадыгина, О.А.; Горячковская, Ю.М.; Яковлев, В.И.; Касулина, В.В.; ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5542ru Citation requested in the work: Клименко О.И., Кондракова И. А., Мадыгина О. А., Горячковская Ю. М., Яковлев В.И., Касулина В.В. 2025. Анализ кооперативного законодательства Российской Федерации и других стран СНГ. ФАО Законодательное исследование № 119. Рим, ФАО. https://doi. org/10.4060/cd5542ru
842
- FAO Комиссия Кодекс Алиментариус. Руководство по процедуре ​ФАО; ВОЗ; ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7978ru Citation requested in the work: ФАО и ВОЗ. 2026. Комиссия Кодекс Алиментариус. Руководство по процедуре. Тридцать первое издание. Рим. https://doi.org/10.4060/cd7978ru
843
  FAO Положение дел в области продовольствия и сельского хозяйства 2024 ФАО ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd2616ru Citation requested in the work: ФАО. 2024. Положение дел в области продовольствия и сельского хозяйства – 2024. Преобразование агропродовольственных систем с ориентацией на ценностные параметры. Рим. https://doi.org/10.4060/cd2616ru
844
  FAO Положение дел в области продовольствия и сельского хозяйства 2025 FAO ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7067ru Citation requested in the work: ФАО. 2025. Положение дел в области продовольствия и сельского хозяйства – 2025. Решение проблемы деградации почв с учетом масштабов землевладений. Рим. https://doi.org/10.4060/cd7067ru
845
  FAO Положение дел на рынках сельскохозяйственной продукции – 2024 FAO ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd2144ru Citation requested in the work: ФАО. 2024. Положение дел на рынках сельскохозяйственной продукции – 2024. Торговля и питание: согласованность политики в интересах обеспечения здорового рациона. Рим. https://doi.org/10.4060/cd2144ru
@@ -848,12 +850,12 @@ FAO Состояние мировых земельных и водных рес
848
  FAO 全球黑土现状 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc3124zh Citation requested in the work: 。中国北京,中国农业出版社。https://doi.org/10.4060/ 粮农组织。2025。《全球黑土现状》 cc3124zh 20-CP 语学院安排并对翻译的准确性及质量负全部责任。如有出入,应以英文原版为准。
849
  FAO 兽药残留对肠道微生物组和人体健康的影响 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5301zh Citation requested in the work: 粮农组织。2025。《兽药残留对肠道微生物组和人体健康的影响——食品安全视角》。中国 北京,中国农业出版社。https://doi.org/10.4060/cc5301zh 20-CP
850
  FAO 农药残留对肠道微生物组和人体健康的影响 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5306zh Citation requested in the work: 粮农组织。2025。《农药残留对肠道微生物组和人体健康的影响——食品安全视角》。中国 北京,中国农业出版社。https://doi.org/10.4060/cc5306zh 20-CP
851
- FAO 卓越数字农业报告 ​FAO; ITU zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc4764zh Citation requested in the work: 粮 农 组 织 和 国 际 电 信 联 盟。2025。《 卓 越 数 字 农 业 报 告 —— 粮 农 组 织 与 国 际 电 联 促 进 欧洲和中亚数字农业良好做法区域竞赛》。中国北京,中国农业出版社。https://doi.org/ 10.4060/cc4764zh 20-CPP2021
852
  FAO 塑料挑战徽章训练手册 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd0922zh Citation requested in the work: 粮农组织。2026。《塑料挑战徽章训练手册》。青年与联合国全球联盟学习和行动系列—— 挑战徽章⑬。中国北京,中国农业出版社。https://doi.org/10.4060/cd0922zh
853
  FAO 满足味蕾的养殖水产品 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5140zh Citation requested in the work: 《满足味蕾的养殖水产品——探索十二种地中海与黑海鱼类从海洋至餐桌之 粮农组织。2025。 旅》。中国北京,中国农业出版社。https://doi.org/10.4060/cc5140zh
854
  FAO 盐碱土探秘 FAO; IUSS; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc0530zh Citation requested in the work: 粮农组织和国际土壤科学联合会。2025。《盐碱土探秘——全球精选十大儿童科普故事》。 中国北京,中国农业出版社。https://doi.org/10.4060/cc0530zh
855
  FAO 社会保护与前瞻行动——保护农业生计 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc7628zh Citation requested in the work: 粮农组织。2026。《社会保护与前瞻行动——保护农业生计》。中国北京,中国农业出版社。 https://doi.org/10.4060/cc7628zh 20-CP 语学院安排并对翻译的准确性及质量负全部责任。如有出入,应以英文原版为准。
856
- FAO 细胞基食品食用安全解析 ​ 粮农组织; 世界卫生组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc4855zh Citation requested in the work: 粮农组织和世卫组织。2025。《细胞基食品食用安全解析》 。中国北京,中国农业出版社。 https://doi.org/10.4060/cc4855zh 20-CP。
857
  FAO 能源挑战徽章:生物能源补充训练手册 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd1397zh Citation requested in the work: 粮农组织。2026。《能源挑战徽章:生物能源补充训练手册》。青年与联合国全球联盟学 习和行动系列。中国北京,中国农业出版社。https://doi.org/10.4060/cd1397zh
858
  FAO 食品法典委员会程序手册 FAO; WHO; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7978zh Citation requested in the work: 粮农组织和世卫组织。2026。《食品法典委员会程序手册》。第三十一版。罗马。https://doi.org/10.4060/cd4216zh
859
  FAO 鱼:知之,烹之,食之 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc1395zh Citation requested in the work: 粮农组织。2025。 《鱼:知之,烹之,食之》。中国北京,中国农业出版社。https://doi.org/10.4060/ cc1395zh
@@ -933,7 +935,7 @@ Kanripo 周禮疑義擧要 江永 (清) zh CC BY-SA 4.0 https://github.com/kanri
933
  Kanripo 嘉靖以來首輔傳 王世貞 (明) zh CC BY-SA 4.0 https://github.com/kanripo/KR2g0037
934
  Kanripo 埤雅 陸佃 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR1j0011
935
  Kanripo 天經惑問 游藝 (清) zh CC BY-SA 4.0 https://github.com/kanripo/KR3f0023
936
- Kanripo 尚書(正文) zh CC BY-SA 4.0 https://github.com/kanripo/KR1b0001
937
  Kanripo 尚書全解 林之奇 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR1b0007
938
  Kanripo 尚書疑義 馬明衡 (明) zh CC BY-SA 4.0 https://github.com/kanripo/KR1b0039
939
  Kanripo 折獄龜鑑 鄭克 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR3c0007
@@ -1028,13 +1030,13 @@ NDLA NDLA: Yrkesfaglig fordypning (EL-ELE vg1) Albertine Aaberge; Bjørn Dølvin
1028
  NDLA NDLA: Yrkesliv i barne- og ungdomsarbeiderfag (HS-BUA vg2) Bente Elisabeth Vetland; Camilla Øvstebø ; Cathrine Dunker Furuly; Einar Martin Kålen; Gro Nedberg Grønlid; Guri Bente Hårberg; Hege Nikolaisen; Karl Henrik Aanesen; Kristin Aase; Kristin Sundstrøm; Riborg Anna Ringereide; Rita Enstad-Karlsen, Terranova Media; Siv Stai; Siv Stai ; Tove Engesvik; Vig no CC BY-SA 4.0 https://ndla.no/subject:1:03e810db-3560-47b5-a5f6-e7afe1d0a2d6
1029
  NDLA NDLA: Yrkesliv i helsearbeiderfag (HS-HEA vg2) Albertine Aaberge; Aleksandra Krogh; Birgit Flaten; Einar Martin Kålen; Hege Nikolaisen; Hege Nikolaisen ; Helene Grotle; Ingrid Schiefloe Myhre; Johannes Leiknes Nag; Karl Henrik Aanesen; Kathrine Synnøve Karlsen; Lars Sandlie/Høgskolen i Lillehammer; Lene Fossbråten; Marit Smith Sørhøy; NAKU; Odd no CC BY-SA 4.0 https://ndla.no/subject:1:f644f829-4e7a-4e74-a63a-342ef786f68a
1030
  OAPEN 100 Cartas para Paulo Freire de quienes pretendemos Enseñar Gárate Vergara, Francisco es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51168
1031
- OAPEN 20 år med fysikkprestasjoner i fritt fall: Analyser fra TIMSS Advanced og andre internasjonale studier Hole, Arne; Onstad, Torgeir; Hagen, Tor Espen no CC BY (OAPEN per-file license code CC-BY) http://library.oapen.org/handle/20.500.12657/23254
1032
  OAPEN Ai margini del contado: Terra, signoria ed élites locali a Sabbion e nel territorio di Cologna Veneta (secoli XII-XIII) STELLA, Attilio it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60437
1033
  OAPEN Bekymringsarbeidet: Politiets forebygging av radikalisering og voldelig ekstremisme Førde, Kristin Engh; Andersen, Arnfinn Jomar; Moum Hellevik, Per no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63675 Citation requested in the work: Førde, K. E, Andersen, A. J. & Hellevik, P. M. (2023). Bekymringsarbeidet. Politiets forebygging av radikalisering og voldelig ekstremisme. Cappelen Damm Akademisk. https://doi.org/10.23865/noasp.185
1034
  OAPEN Bewältigung des Scheiterns: Autobiographische Schriften früherer Parteifunktionäre von NSDAP und SED Danner, Hans-Ulrich de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/91039
1035
  OAPEN Bologna dopo la pandemia: Impatto territoriale e scenari futuri Castrignanò, Marco; RIMONDI, TOMMASO it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60518
1036
  OAPEN Construyendo espacios: la ciudad iberoamericana virreinal: Teoría y estudios de caso Paniagua Pérez, Jesús; Arciello, Daniele es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/51907
1037
- OAPEN Corsi universitari Fortini, Franco it CC BY (OAPEN per-file license code CC-BY) https://library.oapen.org/handle/20.500.12657/89263
1038
  OAPEN Das kolonisierte Heiligtum: Diskriminierungskritische Perspektiven auf das Verfahren der Musealisierung Balzar, Christoph de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://library.oapen.org/handle/20.500.12657/60744
1039
  OAPEN Das sogenannte ‚Königliche Gerichtsbuch‘ – Aufzeichnungen des Michael von Pfullendorf zu den Anfängen des Kammergerichts am römisch-deutschen Königshof (1442 bis 1451): Einführung und Edition Luger, Daniel de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/53427
1040
  OAPEN Del Palacio Negro a la Selva Lacandona: Louis Althusser en México Ortega Reyna, Jaime es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63196
@@ -1045,27 +1047,27 @@ OAPEN Documentación digital y léxico en la traducción e interpretación en lo
1045
  OAPEN En contra de los impíos. La Actuación de la Buena Prensa Católica en la Arquidiócesis de Santiago, 1906-1936 Loyola, Manuel es CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0/) http://library.oapen.org/handle/20.500.12657/32093
1046
  OAPEN Esercizi di ricerca: Dottorato e politiche per la formazione Boffo, Vanna; Togni, Fabio it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62865
1047
  OAPEN Estudios Interculturales desde el Sur: procesos, debates y propuestas Samaniego, Mario es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/50252
1048
- OAPEN Europa: um projecto em construção: Homenagem a David Sassoli Graziani, Michela; Rita, Annabela pt CC BY (OAPEN per-file license code CC-BY) https://library.oapen.org/handle/20.500.12657/62866
1049
  OAPEN Fabulations nocturnes: Écologie, vitalité et opacité dans le cinéma d’Apichatpong Weerasethakul Bordeleau, Érik; Pape, Toni; Rose-Antoinette, Ronald; Szymanski, Adam fr CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) http://library.oapen.org/handle/20.500.12657/31350
1050
  OAPEN Fallbuch Asylrecht: Mit Bezügen zum Aufenthaltsrecht Mantel, Johanna; Nachtigall, Rhea; Wasnick, Lars de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0) https://library.oapen.org/handle/20.500.12657/63515
1051
  OAPEN Firenze prima degli Uberti: Il ceto dirigente fiorentino nell'XI secolo fra riforme diocesane e affermazione personale e familiare Contessa, Maria Pia it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62867
1052
- OAPEN Francesco da Barberino al crocevia: Culture, società, bilinguismo Bischetti, Sara; Montefusco, Antonio it CC BY (OAPEN per-file license code CC-BY) https://library.oapen.org/handle/20.500.12657/52298
1053
  OAPEN Fuentes para una Constitución con Poder Indígena Valenzuela, Esteban; Romero, Natacha es CC BY 2.0 (https://creativecommons.org/licenses/by/2.0/) http://library.oapen.org/handle/20.500.12657/32033
1054
  OAPEN Förändring och kontinuitet: Reformer inom högre utbildning och forskning 1940–2020 Gribbe, Johan sv CC BY 4.0 (http://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/59842 Citation requested in the work: [Johan Gribbe, 2022, Förändring och kontinuitet: Reformer inom högre utbildning och forskning 1940–2020] Universitetskanslersämbetet. DOI: https://doi.org/10.53340/UKAP-4. Licens: CC-BY 4.0
1055
  OAPEN Führt Moral unumgänglich zur Religion?: Zur Kritik der Kantischen Religionsphilosophie bei Jürgen Habermas – eine Entgegnung Langthaler, Rudolf de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51766
1056
  OAPEN Geplante Obsoleszenz: Hinter den Kulissen der Produktentwicklung Poppe, Erik; Longmuß, Jörg de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24340
1057
- OAPEN Gli altri noi: Rom e residenti nella Svizzera italiana: etnografia a mediazione Bizzini, Nadia it CC BY (OAPEN per-file license code CC-BY) https://library.oapen.org/handle/20.500.12657/41430
1058
  OAPEN Gouvernance du secteur de la Sécurité: Leçons des expériences ouest-africaines Bryden, Alan; Chappuis, Fairlie fr CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/32914 Citation requested in the work: Bryden, A et Chappuis, F (dir. publ.) 2015 Gouvernance du secteur de la Sécurité : Leçons des expériences ouest-africaines. London: Ubiquity Press. DOI: http://dx.doi.org/10.5334/bav. Licence: CC-BY 4.0
1059
- OAPEN Guida al mentoring: Aiutare mentori e allievi ad avere successo Chopra, Vineet; Vaughn, Valerie; Saint, Sanjay it CC BY (OAPEN per-file license code CC-BY) https://library.oapen.org/handle/20.500.12657/96163
1060
  OAPEN Guiden till Spaniensverige: Diaspora, integration och transnationalitet bland svenska föreningar i södra Spanien Olsson, Erik sv CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/27488 Citation requested in the work: Olsson, Erik. 2018. Guiden till Spaniensverige: Diaspora, integration och transnationalitet bland svenska föreningar i södra Spanien. Stockholm: Stockholm University Press. DOI: https://doi.org/10.16993/bao. License: CC-BY
1061
  OAPEN Hacia una historia de las tendencias trotskistas después de Trotsky Gaido, Daniel es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58559
1062
  OAPEN Handlungsoptionen auf dem Weg in die Gigabit-Gesellschaft: Eine rechtliche Analyse von Konzessions- und Kooperationsmodellen sowie regulatorischer Entflechtungsbestimmungen Toros, Fabian de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/legalcode) https://library.oapen.org/handle/20.500.12657/50285
1063
  OAPEN Harpe og sverd: Litteraturhistoriske essay om den norske balladen Solberg, Olav no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/50341
1064
  OAPEN Högskolans ansvar: Principer för utveckling av den högre Casson, Andrew sv CC BY 3.0 (https://creativecommons.org/licenses/by/3.0/) http://library.oapen.org/handle/20.500.12657/33044 Citation requested in the work: Casson, A 2015 Högskolans ansvar: Principer för utveckling av den högre utbildningen. London: Ubiquity Press. DOI: http://dx.doi.org/10.5334/bap. License: CC-BY 3.0
1065
- OAPEN Il Fantasma dell’Io. La massa e l’inconscio mimetico: The Phantom of the Ego: Modernism and the Mimetic Unconscious Lawtoo, Nidesh it CC BY (OAPEN per-file license code CC-BY) http://library.oapen.org/handle/20.500.12657/25157
1066
  OAPEN Il video a 360° nella didattica universitaria: Modelli ed esperienze Ranieri, Maria; Luzzi, Damiana; Cuomo, Stefano it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60444
1067
  OAPEN Im Brennpunkt der Wirtschaftspolitik: Innovation, Globalisierung und Klimawandel Keuschnigg, Christian de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/90933
1068
- OAPEN Immaginare l’altrove nell’epoca dell’Antropocene: Media, confini e cambiamenti climatici CAPPI, VALENTINA it CC BY (OAPEN per-file license code CC-BY) https://library.oapen.org/handle/20.500.12657/61650
1069
  OAPEN Inklusionsorientierte Schulentwicklung: Interdisziplinäre Rückblicke, Einblicke und Ausblicke Frohn, Julia; Bengel, Angelika; Piezunka, Anne; Simon, Toni; Dietze, Torsten de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://library.oapen.org/handle/20.500.12657/60549
1070
  OAPEN Innvielse til læreryrket: En analyse av praksislæreres veiledningssamtaler Reier Jensen, Andreas no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24989
1071
  OAPEN Interessekonflikter i forskning Ingierd, Helene; Bay-Larsen, Ingrid; Hiis Hauge, Kjellrun no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/25321
@@ -1088,7 +1090,7 @@ OAPEN La traiettoria storica dell’Etiopia di Meles Zenawi: Fra democrazia rivo
1088
  OAPEN La trama dell’allegoria: Scritture di ricerca e istanza allegorica nel secondo Novecento italiano Caporiccio, Elisa it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58406
1089
  OAPEN La trichera letrada. Intelectuales latinoamericanos y Guerra Fría Alburquerque, Germán es CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0/) http://library.oapen.org/handle/20.500.12657/32089
1090
  OAPEN Les normes de prononciation du français: Une étude perceptive panfrancophone Chalier, Marc fr CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51444
1091
- OAPEN Letras na América Portuguesa: Autores – Textos – Leitores Rodrigues-Moura, Enrique pt CC BY (OAPEN per-file license code CC-BY) https://library.oapen.org/handle/20.500.12657/90323
1092
  OAPEN Lo sguardo territorialista di Leonardo: Il cartografo, l’ingegnere idraulico, il progettista di città e territori Poli, Daniela it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62874
1093
  OAPEN L’intervista immaginata: Da genere mediatico a invenzione letteraria GALLERANI, Guido Mattia it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58403
1094
  OAPEN L’URSS dentro e fuori: La narrazione italiana del mondo sovietico Traini, Cheti it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60443
@@ -1102,7 +1104,7 @@ OAPEN Paradigmas y polifuncionalidad: Estudio diacrónico de «preciso»/«preci
1102
  OAPEN Poéticas espectatoriales en Hispanoamérica y Brasil (1800–1847): Ilustración – emancipación – convivencias excluyentes Fernández, Hans es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/59652
1103
  OAPEN Problemáticas étnicas y sociales desde el pensamiento latinoamericano: Temas, Conceptos, Enfoques Kozel, Andrés; Rawicz, Daniela; Devés, Eduardo es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63199
1104
  OAPEN PROGETTO STREAMING - STRategiE di mitigazione e gestione dei rischi AMbientalI: casi di studio Nel territorio reGionale Toscano: Azioni locali di sostenibilità: cinque progetti per il futuro del territorio toscano Bartalucci, Chiara; Fagioli, Federico; Giachetti, Andrea; NICCOLAI, ALBERTO; Verdi, Leonardo it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58409
1105
- OAPEN Pubblicità, educazione e diritto in Kant Perni, Romina it CC BY (OAPEN per-file license code CC-BY) https://library.oapen.org/handle/20.500.12657/62877
1106
  OAPEN Raccontare la Resistenza a scuola: Esperienze e riflessioni Bravi, Luca; Martinelli, Chiara; Oliviero, Stefano it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60445
1107
  OAPEN Roher Diamant Dalmatien: Die habsburgische Verwaltung, ihre Probleme und das Land, wie beschrieben von seinem Gouverneur Lilienberg fuer Kaiser Franz I. (1834) Clewing, Konrad de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/legalcode) http://library.oapen.org/handle/20.500.12657/26672 Citation requested in the work: Konrad Clewing (Hg.), Roher Diamant Dalmatien. Die habsburgische Verwaltung, ihre Probleme und das Land, wie beschrieben von seinem Gouverneur Lilienberg für Kaiser Franz I. (1834). München, Berlin, Leipzig, Washington/D.C. 2015
1108
  OAPEN Samarbeid om selvhjelp: En antologi om den nye selvhjelpsbevegelsen i Norge Gotaas, Nora; Hatleskog Zeiner, Hilde no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24940
@@ -1114,7 +1116,7 @@ OAPEN Urbane Transformation durch soziale Innovation: Schlüsselbegriffe und Per
1114
  OAPEN Wissenschaftskarrieren und Gender Bias: Chancengerechtigkeit an Hochschulen zwischen formellen Vorgaben und informellen Einflüssen Dahmen-Adkins, Jennifer; Wolffram, Andrea de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/90818
1115
  OAPEN «Parlare di tutto». Un’idea della critica: Il carteggio Baldacci-Fortini Baldacci, Luigi; Fortini, Franco it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62858
1116
  OAPEN Å kjøpe for Norge Langseth, Marius; Similä, Jan Ole no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/49452
1117
- OAPEN Коммуникативный анализ нехудожественного текста для студентов-магистрантов РКИ Perotto, Monica ru CC BY (OAPEN per-file license code CC-BY) https://library.oapen.org/handle/20.500.12657/89272
1118
  OAPEN Конструкции с опорным глаголом в русском и итальянском языках / Support Verb Constructions. A Russian-Italian Contrastive Analysis MAIKO, TATSIANA ru CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60450
1119
  OpenStax Algebra and Trigonometry 2e OpenStax en https://creativecommons.org/licenses/by/4.0/ https://github.com/openstax/osbooks-college-algebra-bundle/blob/4922e46ebc04326979e19392ccc2384cb8b9076c/collections/algebra-and-trigonometry-2e.collection.xml Attribution requested in the work: OpenStax and its content contributors.
1120
  OpenStax American Government 4e OpenStax en https://creativecommons.org/licenses/by/4.0/ https://github.com/openstax/osbooks-american-government/blob/5c90dd7907dbf25f42266417e63ca3bb0012f031/collections/american-government-4e.collection.xml Attribution requested in the work: OpenStax and its content contributors.
 
1
+ # Source-1: per-work credits for the books in the open-books part of its training and validation splits that are
2
+ # not public domain or CC0 (2,846 works). Books and book chapters that came through Common Pile v0.1 (DOAB,
3
+ # Pressbooks, LibreTexts, OER Commons) are credited at collection level in NOTICE 3.5. Each work is credited to
4
+ # the authors and publishers named in it, under the licence recorded for it by the platform it came from (for
5
+ # three works, the IGO licence their own text states; the licence cell says so). Modified: used to train a
6
+ # classifier. No text of these works is distributed with Source-1. This file is part of NOTICE (see NOTICE,
7
+ # part 3).
8
  # citation_or_note: the citation or attribution that the work's own text asks for, as the work gives it (line
9
  # breaks joined, PDF spacing repaired), for FAO books, OpenStax textbooks, Eurydice, JRC and other EU reports, and
10
  # some other books and reports (university-press books, Frontiers ebooks, research and project reports); or a
 
109
  DOAB Questioni di donne: Diplomazia informale e reti femminili alla corte dei Savoia-Carignano (XVII secolo) Lurgo, Elisabetta it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://directory.doabooks.org/handle/20.500.12854/134564
110
  DOAB Religion og etikk i skole og barnehage Afset, Bente; Redse, Arne no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://directory.doabooks.org/handle/20.500.12854/38379
111
  DOAB Réinventer l’art sacré: Le Groupe de Saint-Luc (1919-1945) Noverraz, Camille fr CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://directory.doabooks.org/handle/20.500.12854/152369
112
+ DOAB Sjezd českých právníků 2022 Jednota českých právníků (collective of authors, no editors recorded) cs CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://directory.doabooks.org/handle/20.500.12854/98695
113
  DOAB Slovanský literární svět: kontexty a konfrontace III: Motiv domova ve slovanských literaturách Bujnáková, Jana; Cepková Feješová, Zuzana; Derková, Vladimíra; Eniko, Mateja; Heinigová, Lenka; Hrancová, Hana cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80840
114
  DOAB Somatopedické simulační techniky a intervence: Metodické texty k projektu MUNI 4.0. Pedagogická fakulta, studijní program Logopedie (Bc.) Opatřilová, Dagmar cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80845
115
  DOAB Speciálněpedagogická diagnostika somatopedická: Metodické texty k projektu MUNI 4.0. Pedagogická fakulta, studijní program Logopedie (Bc.) Opatřilová, Dagmar cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80846
 
769
  FAO Commission générale des pêches pour la Méditerranée: Rapport de la quarante-huitième session FAO; fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd9408fr Citation requested in the work: FAO. 2026. Commission générale des pêches pour la Méditerranée: Rapport de la quarante-huitième session – Malaga, Espagne, 4-9 novembre 2025. Commission générale des pêches pour la Méditerranée (CGPM) – Rapports de session, n°48. Rome. https://doi.org/10.4060/cd9408fr
770
  FAO Compendium of decisions made by the Parties to the FAO Agreement on Port State Measures FAO; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0286en Citation requested in the work: FAO. 2026. Compendium of decisions made by the Parties to the FAO Agreement on Port State Measures. Second edition. Rome. https://doi.org/10.4060/ce0286en
771
  FAO Comptes rendus du deuxième Sommet sur la pêche artisanale FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3997fr Citation requested in the work: FAO. 2025. Comptes rendus du deuxième Sommet sur la pêche artisanale, 5-7 juillet 2024, Rome. FAO Comptes rendus des pêches et de l’aquaculture, n° 70. Rome. https://doi.org/10.4060/cd3997fr
772
+ FAO Cостояние мирового рыболовства и аквакультуры – 2026 ФАО; ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd8357ru Citation requested in the work: "ФАО. 2026. Cостояние мирового рыболовства и аквакультуры – 2026. ""Голубая трансформация"": от замысла к практическим результатам. Рим. https://doi.org/10.4060/cd8357ru"
773
  FAO Design of a climate-proof fish buying station Josupeit, H.; Moretti, S.; Sciortino, J.A.; Urbani, R.; van Anrooy, R.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1081en Citation requested in the work: Josupeit, H., Moretti, S., Sciortino, J.A., Urbani, R. & Van Anrooy, R. 2026. Design of a climate-proof fish buying station. FAO Fisheries and Aquaculture Technical Paper, No. 691. Rome, FAO. https://doi.org/10.4060/ce1081en
774
  FAO Developing holistic nutrition guidelines and standards for school meals FAO; WFP; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1121en Citation requested in the work: FAO and WFP. 2026. Developing holistic nutrition guidelines and standards for school meals – A global methodology. Rome. https://doi.org/10.4060/ce1121en
775
  FAO Developing nutrition-sensitive value chains Andrianarimanana, M.; Galante, A.; Liu, B.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0611en Citation requested in the work: Andrianarimanana, M., Galante, A. & Liu, B. 2026. Developing nutrition‑sensitive value chains – Guidelines for practitioners. Rome, FAO. https://doi.org/10.4060/ce0611en
 
820
  FAO Scaling up community-base fisheries management in the Pacific Govan, H.; Tuxson, T.; Tauati, M.; Lalavanua, W.; Schwarz, A.; Kinch, J.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1087en Citation requested in the work: Govan, H., Tuxson,T., Tauati, M., Lalavanua, W., Schwarz, A. & Kinch, J. 2026. Scaling up community-based fisheries management in the Pacific – Outlook and prospects for securing sustainable coastal fisheries, livelihoods and ecosystems. Rome, FAO; Noumea, Pacific Community. https://doi.org/10.4060/ce1087en
821
  FAO Seed to sip: Arabica coffee production guide for Saudi Arabia Gichimu, B.M.; Ghosh, K.; Alfaifi, B.H.; Alfaifi, K.A.; Al Mutlaq, A.M.; Bustamante Adum, D.; Lubabali, H.A.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0725en Citation requested in the work: Gichimu, B.M, Ghosh, K., Alfaifi, B.H., Alfaifi, K.A., Al Mutlaq, A.M., Bustamante Adum, D. & Lubabali, H.A. 2026. Seed to sip: Arabica coffee production guide for Saudi Arabia. Riyadh, FAO. https://doi.org/10.4060/ce0725en
822
  FAO Shaping agrifood systems legislation Rosenbaum, K.L.; Vidar, M.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1406en Citation requested in the work: Rosenbaum, K.L and Vidar, M. 2026. Shaping agrifood systems legislation – Good practices for legal advisors in drafting laws and supporting related processes. Legal Guide No. 5. Rome, FAO. https://doi.org/10.4060/ce1406en
823
+ FAO Soil health and fertilizer Jafari, A.; Le Cotty, T.; Tefft, J.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd9916en Citation requested in the work: Jafari, A., Le Cotty, T. & Tefft, J. 2026. Soil health and fertilizer – Policy and investment prospects in sub-Saharan Africa. Directions in Investment no. 19. Rome, Paris and Brussels, FAO, Agrinatura and the European Union. https://doi.org/10.4060/cd9916en
824
  FAO Status of the world's soil resources 2026 FAO; ITPS; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0792en Citation requested in the work: FAO and ITPS. 2026. Status of the world's soil resources 2026. Rome, FAO. https://doi.org/10.4060/ce0792en
825
  FAO Strengthening small and medium agroenterprise finance in the Near East and North Africa region Aldredge, H.; Priebe, J.; Zook, D.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0863en Citation requested in the work: Aldredge, H., Priebe, J. & Zook, D. 2026. Strengthening small and medium agroenterprise finance in the Near East and North Africa region: Bridging the finance gap. Rome, FAO. https://doi.org/10.4060/ce0863en
826
  FAO Sổ tay hỏi đáp Phùng, T.V.; Tôn, V.Đ.; Sơn, T.H; Phục, N.N.; vi CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0930vi Citation requested in the work: Phùng, T.V., Tôn, V.Đ, Sơn, T.H. and Phục, N.N. 2026. Sổ tay hỏi đáp về th c h nh t t an to n sinh học v xử lý chất thải trong chăn nuôi lợn quy mô vừa v nhỏ. Hà Nội, FAO.
 
834
  FAO Vodič za odgovornu i racionalnu primenu antibiotika kod ovaca i koza Krnjaić, D.; Savić, B.; Trailović, S.; sr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0604sr Citation requested in the work: Krnjaić, D, Savić, B, Trailović, S. 2026. Vodič za odgovornu i racionalnu primenu antibiotika kod ovaca i koza. Beograd, FAO.
835
  FAO Vodič za odgovornu i racionalnu primenu antibiotika kod pasa i mačaka Krnjaić, D.; Resanović, R,; sr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0991sr Citation requested in the work: Krnjaić, D, Resanović, R, Trailović, S. 2026. Vodič za odgovornu i racionalnu primenu antibiotika kod pasa i mačaka . Beograd, FAO.
836
  FAO Voices of cooperators: A collective journey through the agrifood systems in the Near East and North Africa FAO; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0708en Citation requested in the work: FAO. 2026. Voices of cooperators: A collective journey through the agrifood systems in the Near East and North Africa region. Cairo. https://doi.org/10.4060/ce0708en
837
+ FAO Wood products in the bioeconomy Reck, B.K.; Johnston, C.; Foong, A.; Gupta, A.; Karpov, A.; Holsten, A.; Misselwitz, P.; Keenan, R.J.; Formenton Cardoso, N.; Walter, S.; Bull, L.; Steel, E.A.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0315en Citation requested in the work: Reck, B.K., Johnston, C., Foong, A., Gupta, A., Karpov, A., Holsten, A., Misselwitz, P., Keenan, R.J., Formenton Cardoso, N., Walter, S., Bull, L., and Steel, E.A. 2026. Wood products in the bioeconomy: Scenario-based assessment of the potential for engineered wood products in climate change mitigation. Rome, FAO. DOI https://doi.org/10.4060/ce0315en
838
  FAO Évaluation du projet «Consolidation de la paix dans la zone frontalière du nord-est de la Côte d’Ivoire par le biais d’une gestion participative et planifiée des ressources naturelles» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3911fr Citation requested in the work: FAO. 2025. Évaluation du projet «Consolidation de la paix dans la zone frontalière du nord-est de la Côte d’Ivoire, par le biais d’une gestion participative et planifiée des ressources naturelles» – Code du projet: UNJP/IVC/037/PBF. Série évaluation de projet, 01/2025. Rome. https://doi.org/10.4060/cd3911fr
839
  FAO Évaluation du projet «Promouvoir une production de cacao sans déforestation pour réduire les émissions en Côte d’Ivoire» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3686fr Citation requested in the work: FAO. 2025. Évaluation du projet «Promouvoir une production de cacao sans déforestation pour réduire les émissions en Côte d’Ivoire» Rapport de mi-parcous, code du projet: GCP/IVC/609/GCF. Série évaluation de projet, n° 50/2024. Rome. https://doi.org/10.4060/cd3686fr
840
  FAO Évaluation du projet «Soutien à l’auto-emploi de la jeunesse rurale, vecteur de paix et de cohésion sociale au Mali» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7607fr Citation requested in the work: FAO. 2025. Évaluation du projet «Soutien à l’auto-emploi de la jeunesse rurale, vecteur de paix et de cohésion sociale au Mali» Code du projet: UNJP/MLI/068/PBF. Série évaluation de projet, n. 26/2025. Rome. https://doi.org/10.4060/cd7607fr
841
  FAO Évaluation finale du projet «Revitaliser les agroécosystèmes oasiens à travers une approche durable, intégrée et paysagère dans la région de Drâa Tafilalet» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5845fr Citation requested in the work: FAO. 2025. Évaluation finale du projet «Revitaliser les agroécosystèmes oasiens à travers une approche durable, intégrée et paysagère dans la région de Drâa Tafilalet» - Code du projet: GCP/MOR/046/GFF. Série évaluation de projet, N.° 15/2025. Rome. https://doi.org/10.4060/cd5845fr
842
  FAO Évaluation finale du projet «Élimination des polluants organiques persistants et pesticides obsolètes et renforcement de la gestion rationnelle des pesticides au Cameroun» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5148fr Citation requested in the work: FAO. 2025. Évaluation finale du projet «Élimination des polluants organiques persistants et pesticides obsolètes et renforcement de la gestion rationnelle des pesticides au Cameroun» – Code du projet: GCP/CMR/031/GFF, Identifiant FEM: 4641. Série Évaluation de projet, 12/2025 Rome. https://doi.org/10.4060/cd5148fr
843
+ FAO Анализ кооперативного законодательства Российской Федерации и других стран СНГ Клименко, О.И.; Кондракова, И.А.; Мадыгина, О.А.; Горячковская, Ю.М.; Яковлев, В.И.; Касулина, В.В.; ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5542ru Citation requested in the work: Клименко О.И., Кондракова И. А., Мадыгина О. А., Горячковская Ю. М., Яковлев В.И., Касулина В.В. 2025. Анализ кооперативного законодательства Российской Федерации и других стран СНГ. ФАО Законодательное исследование № 119. Рим, ФАО. https://doi.org/10.4060/cd5542ru
844
+ FAO Комиссия Кодекс Алиментариус. Руководство по процедуре ФАО; ВОЗ; ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7978ru Citation requested in the work: ФАО и ВОЗ. 2026. Комиссия Кодекс Алиментариус. Руководство по процедуре. Тридцать первое издание. Рим. https://doi.org/10.4060/cd7978ru
845
  FAO Положение дел в области продовольствия и сельского хозяйства 2024 ФАО ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd2616ru Citation requested in the work: ФАО. 2024. Положение дел в области продовольствия и сельского хозяйства – 2024. Преобразование агропродовольственных систем с ориентацией на ценностные параметры. Рим. https://doi.org/10.4060/cd2616ru
846
  FAO Положение дел в области продовольствия и сельского хозяйства 2025 FAO ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7067ru Citation requested in the work: ФАО. 2025. Положение дел в области продовольствия и сельского хозяйства – 2025. Решение проблемы деградации почв с учетом масштабов землевладений. Рим. https://doi.org/10.4060/cd7067ru
847
  FAO Положение дел на рынках сельскохозяйственной продукции – 2024 FAO ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd2144ru Citation requested in the work: ФАО. 2024. Положение дел на рынках сельскохозяйственной продукции – 2024. Торговля и питание: согласованность политики в интересах обеспечения здорового рациона. Рим. https://doi.org/10.4060/cd2144ru
 
850
  FAO 全球黑土现状 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc3124zh Citation requested in the work: 。中国北京,中国农业出版社。https://doi.org/10.4060/ 粮农组织。2025。《全球黑土现状》 cc3124zh 20-CP 语学院安排并对翻译的准确性及质量负全部责任。如有出入,应以英文原版为准。
851
  FAO 兽药残留对肠道微生物组和人体健康的影响 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5301zh Citation requested in the work: 粮农组织。2025。《兽药残留对肠道微生物组和人体健康的影响——食品安全视角》。中国 北京,中国农业出版社。https://doi.org/10.4060/cc5301zh 20-CP
852
  FAO 农药残留对肠道微生物组和人体健康的影响 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5306zh Citation requested in the work: 粮农组织。2025。《农药残留对肠道微生物组和人体健康的影响——食品安全视角》。中国 北京,中国农业出版社。https://doi.org/10.4060/cc5306zh 20-CP
853
+ FAO 卓越数字农业报告 FAO; ITU zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc4764zh Citation requested in the work: 粮农组织和国际电信联盟。2025。《卓越数字农业报告——粮农组织与国际电联促进欧洲和中亚数字农业良好做法区域竞赛》。中国北京,中国农业出版社。https://doi.org/10.4060/cc4764zh 20-CPP2021
854
  FAO 塑料挑战徽章训练手册 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd0922zh Citation requested in the work: 粮农组织。2026。《塑料挑战徽章训练手册》。青年与联合国全球联盟学习和行动系列—— 挑战徽章⑬。中国北京,中国农业出版社。https://doi.org/10.4060/cd0922zh
855
  FAO 满足味蕾的养殖水产品 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5140zh Citation requested in the work: 《满足味蕾的养殖水产品——探索十二种地中海与黑海鱼类从海洋至餐桌之 粮农组织。2025。 旅》。中国北京,中国农业出版社。https://doi.org/10.4060/cc5140zh
856
  FAO 盐碱土探秘 FAO; IUSS; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc0530zh Citation requested in the work: 粮农组织和国际土壤科学联合会。2025。《盐碱土探秘——全球精选十大儿童科普故事》。 中国北京,中国农业出版社。https://doi.org/10.4060/cc0530zh
857
  FAO 社会保护与前瞻行动——保护农业生计 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc7628zh Citation requested in the work: 粮农组织。2026。《社会保护与前瞻行动——保护农业生计》。中国北京,中国农业出版社。 https://doi.org/10.4060/cc7628zh 20-CP 语学院安排并对翻译的准确性及质量负全部责任。如有出入,应以英文原版为准。
858
+ FAO 细胞基食品食用安全解析 粮农组织; 世界卫生组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc4855zh Citation requested in the work: 粮农组织和世卫组织。2025。《细胞基食品食用安全解析》 。中国北京,中国农业出版社。 https://doi.org/10.4060/cc4855zh 20-CP。
859
  FAO 能源挑战徽章:生物能源补充训练手册 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd1397zh Citation requested in the work: 粮农组织。2026。《能源挑战徽章:生物能源补充训练手册》。青年与联合国全球联盟学 习和行动系列。中国北京,中国农业出版社。https://doi.org/10.4060/cd1397zh
860
  FAO 食品法典委员会程序手册 FAO; WHO; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7978zh Citation requested in the work: 粮农组织和世卫组织。2026。《食品法典委员会程序手册》。第三十一版。罗马。https://doi.org/10.4060/cd4216zh
861
  FAO 鱼:知之,烹之,食之 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc1395zh Citation requested in the work: 粮农组织。2025。 《鱼:知之,烹之,食之》。中国北京,中国农业出版社。https://doi.org/10.4060/ cc1395zh
 
935
  Kanripo 嘉靖以來首輔傳 王世貞 (明) zh CC BY-SA 4.0 https://github.com/kanripo/KR2g0037
936
  Kanripo 埤雅 陸佃 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR1j0011
937
  Kanripo 天經惑問 游藝 (清) zh CC BY-SA 4.0 https://github.com/kanripo/KR3f0023
938
+ Kanripo 尚書(正文) unknown (traditionally attributed to 孔子, Confucius, as compiler) zh CC BY-SA 4.0 https://github.com/kanripo/KR1b0001
939
  Kanripo 尚書全解 林之奇 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR1b0007
940
  Kanripo 尚書疑義 馬明衡 (明) zh CC BY-SA 4.0 https://github.com/kanripo/KR1b0039
941
  Kanripo 折獄龜鑑 鄭克 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR3c0007
 
1030
  NDLA NDLA: Yrkesliv i barne- og ungdomsarbeiderfag (HS-BUA vg2) Bente Elisabeth Vetland; Camilla Øvstebø ; Cathrine Dunker Furuly; Einar Martin Kålen; Gro Nedberg Grønlid; Guri Bente Hårberg; Hege Nikolaisen; Karl Henrik Aanesen; Kristin Aase; Kristin Sundstrøm; Riborg Anna Ringereide; Rita Enstad-Karlsen, Terranova Media; Siv Stai; Siv Stai ; Tove Engesvik; Vig no CC BY-SA 4.0 https://ndla.no/subject:1:03e810db-3560-47b5-a5f6-e7afe1d0a2d6
1031
  NDLA NDLA: Yrkesliv i helsearbeiderfag (HS-HEA vg2) Albertine Aaberge; Aleksandra Krogh; Birgit Flaten; Einar Martin Kålen; Hege Nikolaisen; Hege Nikolaisen ; Helene Grotle; Ingrid Schiefloe Myhre; Johannes Leiknes Nag; Karl Henrik Aanesen; Kathrine Synnøve Karlsen; Lars Sandlie/Høgskolen i Lillehammer; Lene Fossbråten; Marit Smith Sørhøy; NAKU; Odd no CC BY-SA 4.0 https://ndla.no/subject:1:f644f829-4e7a-4e74-a63a-342ef786f68a
1032
  OAPEN 100 Cartas para Paulo Freire de quienes pretendemos Enseñar Gárate Vergara, Francisco es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51168
1033
+ OAPEN 20 år med fysikkprestasjoner i fritt fall: Analyser fra TIMSS Advanced og andre internasjonale studier Hole, Arne; Onstad, Torgeir; Hagen, Tor Espen no CC BY (version not recorded) http://library.oapen.org/handle/20.500.12657/23254
1034
  OAPEN Ai margini del contado: Terra, signoria ed élites locali a Sabbion e nel territorio di Cologna Veneta (secoli XII-XIII) STELLA, Attilio it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60437
1035
  OAPEN Bekymringsarbeidet: Politiets forebygging av radikalisering og voldelig ekstremisme Førde, Kristin Engh; Andersen, Arnfinn Jomar; Moum Hellevik, Per no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63675 Citation requested in the work: Førde, K. E, Andersen, A. J. & Hellevik, P. M. (2023). Bekymringsarbeidet. Politiets forebygging av radikalisering og voldelig ekstremisme. Cappelen Damm Akademisk. https://doi.org/10.23865/noasp.185
1036
  OAPEN Bewältigung des Scheiterns: Autobiographische Schriften früherer Parteifunktionäre von NSDAP und SED Danner, Hans-Ulrich de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/91039
1037
  OAPEN Bologna dopo la pandemia: Impatto territoriale e scenari futuri Castrignanò, Marco; RIMONDI, TOMMASO it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60518
1038
  OAPEN Construyendo espacios: la ciudad iberoamericana virreinal: Teoría y estudios de caso Paniagua Pérez, Jesús; Arciello, Daniele es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/51907
1039
+ OAPEN Corsi universitari Fortini, Franco it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/89263
1040
  OAPEN Das kolonisierte Heiligtum: Diskriminierungskritische Perspektiven auf das Verfahren der Musealisierung Balzar, Christoph de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://library.oapen.org/handle/20.500.12657/60744
1041
  OAPEN Das sogenannte ‚Königliche Gerichtsbuch‘ – Aufzeichnungen des Michael von Pfullendorf zu den Anfängen des Kammergerichts am römisch-deutschen Königshof (1442 bis 1451): Einführung und Edition Luger, Daniel de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/53427
1042
  OAPEN Del Palacio Negro a la Selva Lacandona: Louis Althusser en México Ortega Reyna, Jaime es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63196
 
1047
  OAPEN En contra de los impíos. La Actuación de la Buena Prensa Católica en la Arquidiócesis de Santiago, 1906-1936 Loyola, Manuel es CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0/) http://library.oapen.org/handle/20.500.12657/32093
1048
  OAPEN Esercizi di ricerca: Dottorato e politiche per la formazione Boffo, Vanna; Togni, Fabio it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62865
1049
  OAPEN Estudios Interculturales desde el Sur: procesos, debates y propuestas Samaniego, Mario es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/50252
1050
+ OAPEN Europa: um projecto em construção: Homenagem a David Sassoli Graziani, Michela; Rita, Annabela pt CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/62866
1051
  OAPEN Fabulations nocturnes: Écologie, vitalité et opacité dans le cinéma d’Apichatpong Weerasethakul Bordeleau, Érik; Pape, Toni; Rose-Antoinette, Ronald; Szymanski, Adam fr CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) http://library.oapen.org/handle/20.500.12657/31350
1052
  OAPEN Fallbuch Asylrecht: Mit Bezügen zum Aufenthaltsrecht Mantel, Johanna; Nachtigall, Rhea; Wasnick, Lars de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0) https://library.oapen.org/handle/20.500.12657/63515
1053
  OAPEN Firenze prima degli Uberti: Il ceto dirigente fiorentino nell'XI secolo fra riforme diocesane e affermazione personale e familiare Contessa, Maria Pia it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62867
1054
+ OAPEN Francesco da Barberino al crocevia: Culture, società, bilinguismo Bischetti, Sara; Montefusco, Antonio it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/52298
1055
  OAPEN Fuentes para una Constitución con Poder Indígena Valenzuela, Esteban; Romero, Natacha es CC BY 2.0 (https://creativecommons.org/licenses/by/2.0/) http://library.oapen.org/handle/20.500.12657/32033
1056
  OAPEN Förändring och kontinuitet: Reformer inom högre utbildning och forskning 1940–2020 Gribbe, Johan sv CC BY 4.0 (http://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/59842 Citation requested in the work: [Johan Gribbe, 2022, Förändring och kontinuitet: Reformer inom högre utbildning och forskning 1940–2020] Universitetskanslersämbetet. DOI: https://doi.org/10.53340/UKAP-4. Licens: CC-BY 4.0
1057
  OAPEN Führt Moral unumgänglich zur Religion?: Zur Kritik der Kantischen Religionsphilosophie bei Jürgen Habermas – eine Entgegnung Langthaler, Rudolf de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51766
1058
  OAPEN Geplante Obsoleszenz: Hinter den Kulissen der Produktentwicklung Poppe, Erik; Longmuß, Jörg de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24340
1059
+ OAPEN Gli altri noi: Rom e residenti nella Svizzera italiana: etnografia a mediazione Bizzini, Nadia it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/41430
1060
  OAPEN Gouvernance du secteur de la Sécurité: Leçons des expériences ouest-africaines Bryden, Alan; Chappuis, Fairlie fr CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/32914 Citation requested in the work: Bryden, A et Chappuis, F (dir. publ.) 2015 Gouvernance du secteur de la Sécurité : Leçons des expériences ouest-africaines. London: Ubiquity Press. DOI: http://dx.doi.org/10.5334/bav. Licence: CC-BY 4.0
1061
+ OAPEN Guida al mentoring: Aiutare mentori e allievi ad avere successo Chopra, Vineet; Vaughn, Valerie; Saint, Sanjay it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/96163
1062
  OAPEN Guiden till Spaniensverige: Diaspora, integration och transnationalitet bland svenska föreningar i södra Spanien Olsson, Erik sv CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/27488 Citation requested in the work: Olsson, Erik. 2018. Guiden till Spaniensverige: Diaspora, integration och transnationalitet bland svenska föreningar i södra Spanien. Stockholm: Stockholm University Press. DOI: https://doi.org/10.16993/bao. License: CC-BY
1063
  OAPEN Hacia una historia de las tendencias trotskistas después de Trotsky Gaido, Daniel es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58559
1064
  OAPEN Handlungsoptionen auf dem Weg in die Gigabit-Gesellschaft: Eine rechtliche Analyse von Konzessions- und Kooperationsmodellen sowie regulatorischer Entflechtungsbestimmungen Toros, Fabian de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/legalcode) https://library.oapen.org/handle/20.500.12657/50285
1065
  OAPEN Harpe og sverd: Litteraturhistoriske essay om den norske balladen Solberg, Olav no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/50341
1066
  OAPEN Högskolans ansvar: Principer för utveckling av den högre Casson, Andrew sv CC BY 3.0 (https://creativecommons.org/licenses/by/3.0/) http://library.oapen.org/handle/20.500.12657/33044 Citation requested in the work: Casson, A 2015 Högskolans ansvar: Principer för utveckling av den högre utbildningen. London: Ubiquity Press. DOI: http://dx.doi.org/10.5334/bap. License: CC-BY 3.0
1067
+ OAPEN Il Fantasma dell’Io. La massa e l’inconscio mimetico: The Phantom of the Ego: Modernism and the Mimetic Unconscious Lawtoo, Nidesh it CC BY (version not recorded) http://library.oapen.org/handle/20.500.12657/25157
1068
  OAPEN Il video a 360° nella didattica universitaria: Modelli ed esperienze Ranieri, Maria; Luzzi, Damiana; Cuomo, Stefano it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60444
1069
  OAPEN Im Brennpunkt der Wirtschaftspolitik: Innovation, Globalisierung und Klimawandel Keuschnigg, Christian de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/90933
1070
+ OAPEN Immaginare l’altrove nell’epoca dell’Antropocene: Media, confini e cambiamenti climatici CAPPI, VALENTINA it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/61650
1071
  OAPEN Inklusionsorientierte Schulentwicklung: Interdisziplinäre Rückblicke, Einblicke und Ausblicke Frohn, Julia; Bengel, Angelika; Piezunka, Anne; Simon, Toni; Dietze, Torsten de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://library.oapen.org/handle/20.500.12657/60549
1072
  OAPEN Innvielse til læreryrket: En analyse av praksislæreres veiledningssamtaler Reier Jensen, Andreas no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24989
1073
  OAPEN Interessekonflikter i forskning Ingierd, Helene; Bay-Larsen, Ingrid; Hiis Hauge, Kjellrun no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/25321
 
1090
  OAPEN La trama dell’allegoria: Scritture di ricerca e istanza allegorica nel secondo Novecento italiano Caporiccio, Elisa it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58406
1091
  OAPEN La trichera letrada. Intelectuales latinoamericanos y Guerra Fría Alburquerque, Germán es CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0/) http://library.oapen.org/handle/20.500.12657/32089
1092
  OAPEN Les normes de prononciation du français: Une étude perceptive panfrancophone Chalier, Marc fr CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51444
1093
+ OAPEN Letras na América Portuguesa: Autores – Textos – Leitores Rodrigues-Moura, Enrique pt CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/90323
1094
  OAPEN Lo sguardo territorialista di Leonardo: Il cartografo, l’ingegnere idraulico, il progettista di città e territori Poli, Daniela it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62874
1095
  OAPEN L’intervista immaginata: Da genere mediatico a invenzione letteraria GALLERANI, Guido Mattia it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58403
1096
  OAPEN L’URSS dentro e fuori: La narrazione italiana del mondo sovietico Traini, Cheti it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60443
 
1104
  OAPEN Poéticas espectatoriales en Hispanoamérica y Brasil (1800–1847): Ilustración – emancipación – convivencias excluyentes Fernández, Hans es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/59652
1105
  OAPEN Problemáticas étnicas y sociales desde el pensamiento latinoamericano: Temas, Conceptos, Enfoques Kozel, Andrés; Rawicz, Daniela; Devés, Eduardo es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63199
1106
  OAPEN PROGETTO STREAMING - STRategiE di mitigazione e gestione dei rischi AMbientalI: casi di studio Nel territorio reGionale Toscano: Azioni locali di sostenibilità: cinque progetti per il futuro del territorio toscano Bartalucci, Chiara; Fagioli, Federico; Giachetti, Andrea; NICCOLAI, ALBERTO; Verdi, Leonardo it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58409
1107
+ OAPEN Pubblicità, educazione e diritto in Kant Perni, Romina it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/62877
1108
  OAPEN Raccontare la Resistenza a scuola: Esperienze e riflessioni Bravi, Luca; Martinelli, Chiara; Oliviero, Stefano it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60445
1109
  OAPEN Roher Diamant Dalmatien: Die habsburgische Verwaltung, ihre Probleme und das Land, wie beschrieben von seinem Gouverneur Lilienberg fuer Kaiser Franz I. (1834) Clewing, Konrad de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/legalcode) http://library.oapen.org/handle/20.500.12657/26672 Citation requested in the work: Konrad Clewing (Hg.), Roher Diamant Dalmatien. Die habsburgische Verwaltung, ihre Probleme und das Land, wie beschrieben von seinem Gouverneur Lilienberg für Kaiser Franz I. (1834). München, Berlin, Leipzig, Washington/D.C. 2015
1110
  OAPEN Samarbeid om selvhjelp: En antologi om den nye selvhjelpsbevegelsen i Norge Gotaas, Nora; Hatleskog Zeiner, Hilde no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24940
 
1116
  OAPEN Wissenschaftskarrieren und Gender Bias: Chancengerechtigkeit an Hochschulen zwischen formellen Vorgaben und informellen Einflüssen Dahmen-Adkins, Jennifer; Wolffram, Andrea de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/90818
1117
  OAPEN «Parlare di tutto». Un’idea della critica: Il carteggio Baldacci-Fortini Baldacci, Luigi; Fortini, Franco it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62858
1118
  OAPEN Å kjøpe for Norge Langseth, Marius; Similä, Jan Ole no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/49452
1119
+ OAPEN Коммуникативный анализ нехудожественного текста для студентов-магистрантов РКИ Perotto, Monica ru CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/89272
1120
  OAPEN Конструкции с опорным глаголом в русском и итальянском языках / Support Verb Constructions. A Russian-Italian Contrastive Analysis MAIKO, TATSIANA ru CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60450
1121
  OpenStax Algebra and Trigonometry 2e OpenStax en https://creativecommons.org/licenses/by/4.0/ https://github.com/openstax/osbooks-college-algebra-bundle/blob/4922e46ebc04326979e19392ccc2384cb8b9076c/collections/algebra-and-trigonometry-2e.collection.xml Attribution requested in the work: OpenStax and its content contributors.
1122
  OpenStax American Government 4e OpenStax en https://creativecommons.org/licenses/by/4.0/ https://github.com/openstax/osbooks-american-government/blob/5c90dd7907dbf25f42266417e63ca3bb0012f031/collections/american-government-4e.collection.xml Attribution requested in the work: OpenStax and its content contributors.
EVALUATION.md CHANGED
@@ -1,8 +1,9 @@
1
  # Source-1 evaluation
2
 
3
  Source-1 was compared with its teacher and 16 public quality scorers on three test sets. An independent proprietary
4
- LLM grader scored every chunk with Source-1's 13-field rubric. Each number says how closely a model's ranking agrees
5
- with the grader's, so it measures agreement with Source-1's own rubric, on home ground.
 
6
 
7
  ## Summary
8
 
@@ -38,7 +39,10 @@ were never trained on.
38
 
39
  **Home ground.** The held-out documents come from the same kinds of sources as the training data. The public scorers
40
  were built for their own definitions of quality, most for educational value, and are not wrong when they disagree with
41
- this rubric.
 
 
 
42
 
43
  **The three test sets.**
44
 
@@ -46,7 +50,7 @@ this rubric.
46
  |---|---|---|---|---|
47
  | held-out set (main result) | 495, one per document | 53 | 64 | documents from Source-1's held-out test split, never trained or calibrated on |
48
  | English exam | 414, from 332 documents (413 scored by the teacher) | English | 25 | 200 chunks drawn at random, plus 214 harder cases added on purpose |
49
- | 12-language exam | 352, one per document | 12 | 46 | web text drawn at random from FineWeb-2, about 30 chunks per language |
50
 
51
  **95% intervals.** Each difference between two models comes with a 95% interval from a paired bootstrap over
52
  documents. When the interval excludes zero, chance alone is an unlikely explanation for the difference.
@@ -67,6 +71,8 @@ grammar its card recommends; the serving engine mainly affects speed, so we clai
67
  publishes one.
68
  - The encoder classifiers ran in Hugging Face transformers as their model cards show, in the precision each card
69
  states (bfloat16 where it says so, float32 otherwise).
 
 
70
  - The two fastText models ran in the fasttext library. They read the whole text with newlines turned into spaces, as
71
  the DCLM code does.
72
  - EAI-Distill ran in transformers in float32 with greedy decoding (its repository's default settings sample).
@@ -132,10 +138,12 @@ What grades from the grader's model family did and did not influence:
132
  - **Held-out set** (the main result): 495 chunks in 53 languages, one per document. All come from Source-1's held-out
133
  test split, drawn from the four data stages (web 129, multilingual web 276, conversations/code/synthetic 48, open
134
  books 42). The grader drops 64 of them. 159 chunks are English and 29 Chinese; most other languages have 1 to 12
135
- chunks. Source-1 never trained or calibrated on these documents. They come from the same kinds of sources as the
136
- training data. 42 of the 495 chunks (8.5%) come from sources that were later removed from training under the
137
- license rules (19 from DCLM-baseline, 16 raw Common Crawl pages, and 7 from collections with unreliable or gated
138
- license terms). The test split also keeps documents that the license filtering removed from training.
 
 
139
  - **English exam**: 414 chunks from 332 whole documents split into chunks. 200 chunks from 196 documents were drawn at
140
  random from web, wiki, Common Pile, math and code sources (the **random-sample** chunks and documents). 214 chunks
141
  from 136 documents were added on purpose to cover harder cases (long documents 57 chunks, academic 50, math 28,
@@ -145,12 +153,14 @@ What grades from the grader's model family did and did not influence:
145
  one of the 414 chunks (a random-sample chunk). So the comparisons with the teacher and the public scorers' English
146
  rows use the 413 chunks it scored (411 or 412 where a public scorer also lacks one). The 512-token comparison uses
147
  all 414.
148
- - **12-language exam**: 352 chunks, one per document, all drawn at random from FineWeb-2 web text in 12 languages
149
- (ar, bn, de, es, hi, ja, ko, ru, sw, th, vi, zh; about 30 each). It has no code or math: 351 of the 352 chunks are
150
- plain text by the grader's label. The grader drops 46.
 
151
 
152
  Every exam document was checked against the training data (exact text, URL, title, long-line and shingle matching).
153
- No copies were found.
 
154
 
155
  #### Near-duplicate check
156
 
@@ -162,9 +172,9 @@ each, all short templated texts. 9 in all are covered 20% or more. These chunks
162
 
163
  #### Safety filter
164
 
165
- A safety filter fixed before any result was seen removed a small number of documents from every evaluation set and
166
- from Source-1's training, validation and test data. Documents that substantially copy removed text were removed from
167
- its data as well. Every model is compared on the same filtered chunks.
168
 
169
  </details>
170
 
@@ -178,6 +188,10 @@ This file calls either "the grader". That model family also includes one of the
178
  rubric text. The exams were graded with an older revision of the rubric, from before the rubric settled how to score
179
  ads.
180
 
 
 
 
 
181
  The grader returns the 13 fields only. Its reference overall score and keep flag are computed from its scores in the
182
  same way as Source-1's: the overall formula of the rubric (see [README.md](README.md#what-you-get)), and keep =
183
  false when the rubric's hard filters match (`toxicity >= 4`, `spam_seo >= 4` or `boilerplate >= 4.5`). "The grader's
@@ -188,10 +202,9 @@ under [How the public scorers were run](#how-the-public-scorers-were-run).
188
 
189
  #### Grader self-agreement
190
 
191
- The exam grader also graded part of each exam a second time. The two gradings agree at rank 0.904 on 97 English
192
- random-sample documents and 0.945 on 58 documents of the 12-language exam. This is a rough noise ceiling: it shows how
193
- much one grading moves when repeated. It is not a hard upper bound, because a model can agree with one grading more
194
- closely than two noisy gradings agree with each other.
195
 
196
  </details>
197
 
@@ -252,9 +265,10 @@ the grader keeps ([Metric definitions](#metric-definitions)).
252
  | propella-1 1.7B | 495 | 0.736 | 0.855 | 0.896 | 0.599 (about 38/64) |
253
  | JQL-Edu (mean of 3 balanced heads) | 495 | 0.600 | 0.737 | 0.826 | 0.328 (21/64) |
254
 
255
- - At this matched rate Source-1 catches 46 of the grader's 64 drops and the teacher 44, a difference within noise. At
256
- each model's own drop line (for the teacher, its own keep flags) they catch 43 and 46
257
- ([The drop line](#the-drop-line)). Source-1 and the teacher are also level within noise on rank and AUC.
 
258
  - The lead over propella-1 4B, the closest public scorer, is at least as large outside English: 0.916 against 0.765
259
  on the 335 non-English chunks.
260
 
@@ -530,7 +544,7 @@ At the shipped drop line, counted per document (the teacher at its own keep flag
530
  |---|---|---|---|---|---|
531
  | English, all 332 | 24 | 14 (0.583) | 96.7% | 14 (0.583) | 96.1% |
532
  | English, the 196 random-sample documents | 13 | 9 (0.692) | 97.4% | 6 (0.462) | 94.9% |
533
- | 12 languages, all 352 (all random) | 46 | 31 (0.674) | 94.3% | 28 (0.609) | 93.5% |
534
 
535
  The exams were graded with an older revision of the rubric, before it settled how to score ads (`spam_seo` 3, kept).
536
  So part of the gap is rubric drift that affects the teacher and Source-1 alike.
@@ -707,7 +721,8 @@ the public scorers or the teacher was made under the same conditions, so none is
707
  Precision and batching: computing in bfloat16, both weight files give the same scores. On these 1,261 chunks, scoring
708
  each chunk alone instead of in the default batches moved `overall` by up to 0.04 and a single field by up to 0.10 (3
709
  labels changed, no keep decision). float32 differs from bfloat16 by a similar amount (up to 0.03 on `overall` and 0.09
710
- on a single field; 4 labels and 1 keep decision changed).
 
711
 
712
  </details>
713
 
@@ -716,7 +731,11 @@ on a single field; 4 labels and 1 keep decision changed).
716
  - **Home ground.** It measures agreement with Source-1's rubric, as applied by graders from one proprietary model
717
  family whose grades also steered development. It does not show how Source-1 does on text from other sources, or
718
  that filtering with it trains better language models.
 
 
719
  - **The exams helped choose the teacher**, so they are not independent of the grader.
 
 
720
  - **It misses about a third of the grader's drops** at the shipped line: it catches 43 of 64 on the held-out set, the
721
  teacher, at its own keep flags, 46.
722
  - **Like its teacher, it rates some qualities higher than the grader does.** On the English exam, `educational_value`
@@ -739,7 +758,9 @@ on a single field; 4 labels and 1 keep decision changed).
739
  <summary>The full text of each limitation</summary>
740
 
741
  - **Home-ground evaluation.** The benchmark measures agreement with Source-1's rubric as applied by independent LLM
742
- graders from one proprietary model family (one graded the exams, another the held-out set). Grades from that family,
 
 
743
  among other proprietary LLM graders, also steered the rubric revisions, the choice of the teacher and its prompt
744
  setup, and the drop-line candidates and floor ([Independence from development](#independence-from-development)).
745
  The held-out documents come from the same kinds of sources as the training data. The public scorers were built for
@@ -747,8 +768,15 @@ on a single field; 4 labels and 1 keep decision changed).
747
  more generous reading of propella-1 halves its gap to Source-1 ([How propella-1 is read](#how-propella-1-is-read)).
748
  The comparison does not show how Source-1 does on text from other sources, or that filtering with Source-1 trains
749
  better language models; neither has been tested.
 
 
 
750
  - **The exams are not independent of the grader.** They are the samples on which the teacher and its prompt setup were
751
  chosen against the exam grader's labels. The held-out set, which played no part in that choice, is the main result.
 
 
 
 
752
  - **It misses about a third of the chunks the grader drops at the shipped line.** On the held-out set the shipped line
753
  catches 67.2% of the grader's drops (43/64; interval 55.0% to 77.4%); the teacher catches 71.9% (46/64). 17 of
754
  Source-1's 21 misses are also missed by the teacher, so most of what Source-1 misses its teacher misses too. Most
@@ -842,8 +870,9 @@ cannot be recomputed from this repository alone.
842
  | code_quality | Not usable code: garbled, minified, obfuscated | Very poor: likely non-functional fragments, no structure, or auto-generated boilerplate | Poor: may work but messy | Acceptable: readable, plausibly correct, minimal docs | Good: clean, idiomatic, documented | Excellent: exemplary, production quality, instructive |
843
  | math_quality | Garbled math | Mostly wrong or incoherent | Some correct math, but errors or skipped steps | Generally correct, key steps shown | Correct, clean notation, complete steps | Rigorous and elegant, every step justified |
844
 
845
- The rubric anchors and the head layout are in `source1.json`. The teacher's prompt also had a few special rules that
846
- are not in `source1.json`, and Source-1 was trained on labels that follow them, as far as the teacher did:
 
847
 
848
  - Pages whose main purpose is to promote or sell a business, product or service (company "about us" pages, product
849
  and landing pages, shop listings, brochures) are ads: format `product_page` and `spam_seo` 3, even when cleanly
@@ -931,8 +960,9 @@ filtering removed, for evaluation only):
931
  - Separately, a pre-specified safety filter removed documents from every split, and a rule fixed in advance also
932
  removed documents that substantially copy text the safety filter removed (every split).
933
  - Book pages that are mostly a table of contents were left out of training (99 chunks).
934
- - Every exam document was checked against the training data (exact text, URL, title, long-line and shingle matching);
935
- no copies were found.
 
936
 
937
  #### Labels
938
 
 
1
  # Source-1 evaluation
2
 
3
  Source-1 was compared with its teacher and 16 public quality scorers on three test sets. An independent proprietary
4
+ LLM grader scored every chunk with Source-1's 13-field rubric: its grades were never trained on, and on the held-out
5
+ set it was given the same instructions as Source-1's teacher. Each number says how closely a model's ranking agrees
6
+ with the grader's, so it measures agreement with this grader applying Source-1's own rubric, on home ground.
7
 
8
  ## Summary
9
 
 
39
 
40
  **Home ground.** The held-out documents come from the same kinds of sources as the training data. The public scorers
41
  were built for their own definitions of quality, most for educational value, and are not wrong when they disagree with
42
+ this rubric. Even a scorer that matched the grader's own `educational_value` scores exactly would reach only 0.874 on
43
+ the held-out set, 0.900 on the English exam and 0.817 on the 12-language exam, so a small part of the gap to the
44
+ educational-value classifiers comes from what this measure asks for;
45
+ [Educational value alone](#educational-value-alone) is the fairer comparison for them.
46
 
47
  **The three test sets.**
48
 
 
50
  |---|---|---|---|---|
51
  | held-out set (main result) | 495, one per document | 53 | 64 | documents from Source-1's held-out test split, never trained or calibrated on |
52
  | English exam | 414, from 332 documents (413 scored by the teacher) | English | 25 | 200 chunks drawn at random, plus 214 harder cases added on purpose |
53
+ | 12-language exam | 352, one per document | 12 | 46 | web text drawn from FineWeb-2's test split (blocks of 10 consecutive rows at random positions; Spanish from its first rows), about 30 chunks per language |
54
 
55
  **95% intervals.** Each difference between two models comes with a 95% interval from a paired bootstrap over
56
  documents. When the interval excludes zero, chance alone is an unlikely explanation for the difference.
 
71
  publishes one.
72
  - The encoder classifiers ran in Hugging Face transformers as their model cards show, in the precision each card
73
  states (bfloat16 where it says so, float32 otherwise).
74
+ - Source-1 itself ran in bfloat16, its default. In float32 its rank agreement on the held-out set is unchanged to
75
+ three decimal places.
76
  - The two fastText models ran in the fasttext library. They read the whole text with newlines turned into spaces, as
77
  the DCLM code does.
78
  - EAI-Distill ran in transformers in float32 with greedy decoding (its repository's default settings sample).
 
138
  - **Held-out set** (the main result): 495 chunks in 53 languages, one per document. All come from Source-1's held-out
139
  test split, drawn from the four data stages (web 129, multilingual web 276, conversations/code/synthetic 48, open
140
  books 42). The grader drops 64 of them. 159 chunks are English and 29 Chinese; most other languages have 1 to 12
141
+ chunks. Chinese was oversampled on purpose: 15 of its 29 chunks were added to the proportional sample, and they hold
142
+ 7 of the 64 grader drops. 26 of the 53 languages have 5 or fewer chunks. Source-1 never trained or calibrated on
143
+ these documents. They come from the same kinds of sources as the training data. 42 of the 495 chunks (8.5%) come
144
+ from sources that were later removed from training under the license rules (19 from DCLM-baseline, 16 raw Common
145
+ Crawl pages, and 7 from collections with unreliable or gated license terms). The test split also keeps documents
146
+ that the license filtering removed from training.
147
  - **English exam**: 414 chunks from 332 whole documents split into chunks. 200 chunks from 196 documents were drawn at
148
  random from web, wiki, Common Pile, math and code sources (the **random-sample** chunks and documents). 214 chunks
149
  from 136 documents were added on purpose to cover harder cases (long documents 57 chunks, academic 50, math 28,
 
153
  one of the 414 chunks (a random-sample chunk). So the comparisons with the teacher and the public scorers' English
154
  rows use the 413 chunks it scored (411 or 412 where a public scorer also lacks one). The 512-token comparison uses
155
  all 414.
156
+ - **12-language exam**: 352 chunks, one per document, all FineWeb-2 web text in 12 languages (ar, bn, de, es, hi,
157
+ ja, ko, ru, sw, th, vi, zh; about 30 each), drawn from FineWeb-2's test split (blocks of 10 consecutive rows at
158
+ random positions; Spanish from its first rows). It has no code or math: 351 of the 352 chunks are plain text by the
159
+ grader's label. The grader drops 46.
160
 
161
  Every exam document was checked against the training data (exact text, URL, title, long-line and shingle matching).
162
+ None has half or more of its text in a training document; the largest share of an exam text found in a training
163
+ document is 41%.
164
 
165
  #### Near-duplicate check
166
 
 
172
 
173
  #### Safety filter
174
 
175
+ A safety filter, with a rule fixed before it was first run, removed a small number of documents from every evaluation
176
+ set and from Source-1's training, validation and test data. Documents that substantially copy removed text were
177
+ removed from its data as well. Every model is compared on the same filtered chunks.
178
 
179
  </details>
180
 
 
188
  rubric text. The exams were graded with an older revision of the rubric, from before the rubric settled how to score
189
  ads.
190
 
191
+ The held-out grader was given the teacher's own instructions word for word: the 13-field rubric and the special rules
192
+ in [Appendix A](#appendix-a-rubric-anchors), for example that ads are scored `spam_seo` 3 and kept. The public scorers
193
+ follow none of these rules.
194
+
195
  The grader returns the 13 fields only. Its reference overall score and keep flag are computed from its scores in the
196
  same way as Source-1's: the overall formula of the rubric (see [README.md](README.md#what-you-get)), and keep =
197
  false when the rubric's hard filters match (`toxicity >= 4`, `spam_seo >= 4` or `boilerplate >= 4.5`). "The grader's
 
202
 
203
  #### Grader self-agreement
204
 
205
+ The exam grader also graded part of each exam a second time. The two gradings agree at rank 0.902 on 97 English
206
+ random-sample chunks and 0.945 on 58 chunks of the 12-language exam. On exactly these chunks Source-1 reaches 0.886
207
+ and 0.892 and the teacher 0.879 and 0.882, below the grader's agreement with itself (within noise for English).
 
208
 
209
  </details>
210
 
 
265
  | propella-1 1.7B | 495 | 0.736 | 0.855 | 0.896 | 0.599 (about 38/64) |
266
  | JQL-Edu (mean of 3 balanced heads) | 495 | 0.600 | 0.737 | 0.826 | 0.328 (21/64) |
267
 
268
+ - At this matched rate Source-1 catches 46 of the grader's 64 drops and the teacher 44, a difference within noise
269
+ (without the 15 added Chinese chunks: about 39 and 39 of 57). At each model's own drop line (for the teacher, its
270
+ own keep flags) they catch 43 and 46 ([The drop line](#the-drop-line)). Source-1 and the teacher are also level
271
+ within noise on rank and AUC.
272
  - The lead over propella-1 4B, the closest public scorer, is at least as large outside English: 0.916 against 0.765
273
  on the 335 non-English chunks.
274
 
 
544
  |---|---|---|---|---|---|
545
  | English, all 332 | 24 | 14 (0.583) | 96.7% | 14 (0.583) | 96.1% |
546
  | English, the 196 random-sample documents | 13 | 9 (0.692) | 97.4% | 6 (0.462) | 94.9% |
547
+ | 12 languages, all 352 (none added on purpose) | 46 | 31 (0.674) | 94.3% | 28 (0.609) | 93.5% |
548
 
549
  The exams were graded with an older revision of the rubric, before it settled how to score ads (`spam_seo` 3, kept).
550
  So part of the gap is rubric drift that affects the teacher and Source-1 alike.
 
721
  Precision and batching: computing in bfloat16, both weight files give the same scores. On these 1,261 chunks, scoring
722
  each chunk alone instead of in the default batches moved `overall` by up to 0.04 and a single field by up to 0.10 (3
723
  labels changed, no keep decision). float32 differs from bfloat16 by a similar amount (up to 0.03 on `overall` and 0.09
724
+ on a single field; 4 labels and 1 keep decision changed); computing in float32 with the default bfloat16 weights moved
725
+ `overall` by up to 0.05. In float32, scores do not depend on the batch.
726
 
727
  </details>
728
 
 
731
  - **Home ground.** It measures agreement with Source-1's rubric, as applied by graders from one proprietary model
732
  family whose grades also steered development. It does not show how Source-1 does on text from other sources, or
733
  that filtering with it trains better language models.
734
+ - **The numbers are agreement with one grader family, not accuracy.** Another grader applying the same rubric would
735
+ give different values.
736
  - **The exams helped choose the teacher**, so they are not independent of the grader.
737
+ - **Agreement is lower among good texts.** Among the chunks the grader keeps, Source-1's rank agreement is 0.87, and in
738
+ the better half of those 0.73.
739
  - **It misses about a third of the grader's drops** at the shipped line: it catches 43 of 64 on the held-out set, the
740
  teacher, at its own keep flags, 46.
741
  - **Like its teacher, it rates some qualities higher than the grader does.** On the English exam, `educational_value`
 
758
  <summary>The full text of each limitation</summary>
759
 
760
  - **Home-ground evaluation.** The benchmark measures agreement with Source-1's rubric as applied by independent LLM
761
+ graders from one proprietary model family (one graded the exams, another the held-out set). They are independent in
762
+ that their grades were never trained on; the held-out grader was given the teacher's own instructions word for word
763
+ ([The grader in detail](#the-grader-in-detail)). Grades from that family,
764
  among other proprietary LLM graders, also steered the rubric revisions, the choice of the teacher and its prompt
765
  setup, and the drop-line candidates and floor ([Independence from development](#independence-from-development)).
766
  The held-out documents come from the same kinds of sources as the training data. The public scorers were built for
 
768
  more generous reading of propella-1 halves its gap to Source-1 ([How propella-1 is read](#how-propella-1-is-read)).
769
  The comparison does not show how Source-1 does on text from other sources, or that filtering with Source-1 trains
770
  better language models; neither has been tested.
771
+ - **The numbers are agreement with one grader family, not accuracy.** Another grader applying the same rubric would
772
+ give different values, and differences of a few hundredths near the top (Source-1 against its teacher) may reflect
773
+ this grader's own habits.
774
  - **The exams are not independent of the grader.** They are the samples on which the teacher and its prompt setup were
775
  chosen against the exam grader's labels. The held-out set, which played no part in that choice, is the main result.
776
+ - **Agreement is lower among good texts.** About 13% of the held-out chunks are spam, boilerplate or toxic, which are
777
+ easy to tell apart. Among the chunks the grader keeps, Source-1's rank agreement is 0.87, and in the better half of
778
+ those 0.73 (teacher 0.75, propella-1 4B 0.61). Every scorer drops like this on already-filtered text; if you rank
779
+ filtered text, expect the lower figure.
780
  - **It misses about a third of the chunks the grader drops at the shipped line.** On the held-out set the shipped line
781
  catches 67.2% of the grader's drops (43/64; interval 55.0% to 77.4%); the teacher catches 71.9% (46/64). 17 of
782
  Source-1's 21 misses are also missed by the teacher, so most of what Source-1 misses its teacher misses too. Most
 
870
  | code_quality | Not usable code: garbled, minified, obfuscated | Very poor: likely non-functional fragments, no structure, or auto-generated boilerplate | Poor: may work but messy | Acceptable: readable, plausibly correct, minimal docs | Good: clean, idiomatic, documented | Excellent: exemplary, production quality, instructive |
871
  | math_quality | Garbled math | Mostly wrong or incoherent | Some correct math, but errors or skipped steps | Generally correct, key steps shown | Correct, clean notation, complete steps | Rigorous and elegant, every step justified |
872
 
873
+ The rubric anchors and the head layout are in `source1.json`. The teacher's prompt, which the held-out grader also
874
+ received word for word, had a few special rules that are not in `source1.json`. Source-1 was trained on labels that
875
+ follow them, as far as the teacher did:
876
 
877
  - Pages whose main purpose is to promote or sell a business, product or service (company "about us" pages, product
878
  and landing pages, shop listings, brochures) are ads: format `product_page` and `spam_seo` 3, even when cleanly
 
960
  - Separately, a pre-specified safety filter removed documents from every split, and a rule fixed in advance also
961
  removed documents that substantially copy text the safety filter removed (every split).
962
  - Book pages that are mostly a table of contents were left out of training (99 chunks).
963
+ - Every exam document was checked against the training data (exact text, URL, title, long-line and shingle matching).
964
+ None has half or more of its text in a training document; the largest share of an exam text found in a training
965
+ document is 41%.
966
 
967
  #### Labels
968
 
NOTICE CHANGED
@@ -63,10 +63,11 @@ are not part of this release.
63
 
64
  Source-1 was evaluated with an independent proprietary LLM grader, used for evaluation
65
  only: the grader's outputs were never used as training targets or training data, and
66
- the shipped drop line was chosen by a fixed rule on the teacher's labels. Grades and
67
- reviews by proprietary LLMs from the grader's model family did inform some design
68
- choices (the rubric revision, the choice of teacher and its prompt on the exam sets,
69
- the candidate drop lines and which collections were filtered out); see the
 
70
  development note in README.md and "Independence from development" in EVALUATION.md.
71
 
72
 
@@ -100,13 +101,16 @@ states no open licence; web pages under NC or ND terms, and pages on sites that
100
  other people's documents; and code whose licence or file header is copyleft,
101
  proprietary or not on a permissive allow-list.
102
 
103
- Every book that is not public domain or CC0 is also credited individually in the file
104
- CREDITS_BOOKS.tsv, which is part of this notice: title, authors, language, licence and
105
- source URL. Where a work's own text asks to be cited or attributed in a particular way,
106
- the file gives that citation or attribution: FAO's "Required citation", OpenStax's
107
- attribution request, the citations requested by Eurydice, JRC and other EU reports, and
108
- those of some other books and reports (university presses, Frontiers ebooks, research and
109
- project reports). The World Bank works are credited in Appendix A.
 
 
 
110
 
111
  License links used below:
112
  ODC-By 1.0 https://opendatacommons.org/licenses/by/1-0/
@@ -372,9 +376,12 @@ licence were removed before training.
372
  4. CONTACT AND REMOVAL REQUESTS
373
  ================================================================================
374
 
375
- Questions about these credits, corrections, and removal or opt-out requests:
376
- the Community tab of the model repository,
377
- https://huggingface.co/msmth/Source-1/discussions.
 
 
 
378
 
379
 
380
  ================================================================================
 
63
 
64
  Source-1 was evaluated with an independent proprietary LLM grader, used for evaluation
65
  only: the grader's outputs were never used as training targets or training data, and
66
+ the shipped drop line was chosen by a fixed rule on the teacher's labels. On the
67
+ held-out set the grader was given the same instructions as the teacher. Grades and
68
+ reviews by proprietary LLMs, among them the grader's model family, did inform some
69
+ design choices (the rubric revision, the choice of teacher and its prompt on the exam
70
+ sets, the candidate drop lines and which collections were filtered out); see the
71
  development note in README.md and "Independence from development" in EVALUATION.md.
72
 
73
 
 
101
  other people's documents; and code whose licence or file header is copyleft,
102
  proprietary or not on a permissive allow-list.
103
 
104
+ Every book in the open-books part of the training data that is not public domain or
105
+ CC0 is also credited individually in the file CREDITS_BOOKS.tsv, which is part of this
106
+ notice: title, authors, language, licence and source URL. Where a work's own text asks
107
+ to be cited or attributed in a particular way, the file gives that citation or
108
+ attribution: FAO's "Required citation", OpenStax's attribution request, the citations
109
+ requested by Eurydice, JRC and other EU reports, and those of some other books and
110
+ reports (university presses, Frontiers ebooks, research and project reports). The World
111
+ Bank works are credited in Appendix A. Books and book chapters that came through
112
+ Common Pile v0.1 (DOAB, Pressbooks, LibreTexts, OER Commons) are credited at collection
113
+ level in 3.5.
114
 
115
  License links used below:
116
  ODC-By 1.0 https://opendatacommons.org/licenses/by/1-0/
 
376
  4. CONTACT AND REMOVAL REQUESTS
377
  ================================================================================
378
 
379
+ Questions about these credits, corrections, and removal or opt-out requests: open a
380
+ discussion in the Community tab of the model repository,
381
+ https://huggingface.co/msmth/Source-1/discussions. If your request involves personal
382
+ information, open a discussion without the details and we will arrange a private way to
383
+ reach us. We review every request and, where the content is in our training data,
384
+ exclude it from future versions; published weights cannot be changed.
385
 
386
 
387
  ================================================================================
README.md CHANGED
@@ -146,16 +146,19 @@ languages, up to 8,192 tokens at a time. It has 307M parameters and is free to u
146
  ## Highlights
147
 
148
  The numbers are rank agreement with an independent proprietary LLM grader: how closely a model puts texts in the same
149
- order as the grader does (1.0 means the same order, 0 means no link).
150
-
151
- - **Beats all 16 public quality scorers we tested, including FineWeb-Edu and propella-1, on all three test sets.** On the main test set (495 held-out
152
- texts in 53 languages): 0.90 vs 0.76 for the best of them, propella-1 4B, a model 13x larger.
153
- - **Nearly matches its teacher**, the open-weight 27B LLM that labeled its training data, at 1/88 the size: 0.90 vs 0.91
154
- on the main test set.
 
 
155
  - **Also ahead when judged on educational value alone**, the thing most public scorers were built for.
156
 
157
  *Caveat: the grader scored with Source-1's own rubric (its scoring guide), so this test plays to Source-1's
158
- strengths. See [Limitations](#limitations) and [EVALUATION.md](EVALUATION.md).*
 
159
 
160
  ## Quick start
161
 
@@ -178,6 +181,16 @@ doc = model.score("Photosynthesis is how plants turn light, water and carbon dio
178
  print(doc["overall"], doc["keep"]) # a 0-5 score and the keep/drop decision
179
  ```
180
 
 
 
 
 
 
 
 
 
 
 
181
  <details>
182
  <summary>More usage: many texts, precision, options, speed, files</summary>
183
 
@@ -206,16 +219,22 @@ python source1.py --model . --input page.txt --device cpu --precision fp32 --dty
206
  ```
207
 
208
  - **Output.** One flat dict per document with the 13 fields, `overall`, `keep` and `drop_reasons`. It also gives the
209
- length in tokens, the number of chunks (`parts`), whether a chunk was cut, and per-chunk results (`chunks`).
 
 
210
  - **Precision.** By default it loads `model.safetensors` (bfloat16, 0.6 GB). For the full float32 weights (1.2 GB),
211
  remove `--exclude` from the download and pass `precision="fp32"`. It computes in bfloat16 on NVIDIA Ampere or newer
212
  GPUs and in float32 elsewhere. float16 is not supported. In bfloat16, scores can shift by up to about 0.04
213
  depending on which texts share a batch; pass `dtype="fp32"` or `batch_tokens=1` to avoid that.
214
  - **Options.** `drop_line` sets your own drop rule, such as `"toxicity >= 4 or spam_seo >= 3 or boilerplate >= 4.5"`.
215
  `max_chunks` scores only some chunks of very long documents, for speed. `revision` pins a Hub version. The command
216
- line reads JSONL, JSON or plain text, and `python source1.py --help` lists every flag.
 
217
  - **Examples and speed.** `examples/` holds six documents and their expected CPU output. A GPU can differ by a few
218
  hundredths. Source-1 scores about 30 chunks (50,000 tokens) per second on one RTX 3090 in bfloat16.
 
 
 
219
 
220
  **Files**
221
 
@@ -229,12 +248,12 @@ python source1.py --model . --input page.txt --device cpu --precision fp32 --dty
229
  | `calibration.json` | the drop line and the quality-score offsets, with how they were chosen |
230
  | `tokenizer.json`, `tokenizer_config.json` | mmBERT's tokenizer (same vocabulary and merges, re-saved) |
231
  | `source1.py` | standalone loader, Python API and command line |
232
- | `requirements.txt` | `torch`, `transformers`, `safetensors`, `tokenizers` |
233
  | `examples/` | six sample inputs (`sample.jsonl`) and their expected command-line output (`expected_output.jsonl`) |
234
  | `EVALUATION.md` | the full evaluation, the rubric anchors and the training details |
235
  | `images/` | the benchmark chart above |
236
  | `LICENSE`, `NOTICE`, `AUTHORS` | license text, third-party notices and credits, authors |
237
- | `CREDITS_BOOKS.tsv` | per-work credits for the training books that are not public domain (part of `NOTICE`) |
238
 
239
  </details>
240
 
@@ -260,9 +279,11 @@ Source-1 returns these 13 fields. Higher is better for quality scores and worse
260
  scores are `null` when they do not apply. You also get:
261
 
262
  - `overall`: one 0-5 score, the quality scores minus penalties for red flags.
263
- - `keep`: false (drop) if `toxicity >= 4` or `spam_seo >= 3.5` or `boilerplate >= 4.5`.
 
264
 
265
- Long documents are split into chunks. Each chunk is scored, and the results are combined into one.
 
266
 
267
  <details>
268
  <summary>All fields in detail, the scoring formula, the drop line and long documents</summary>
@@ -277,8 +298,10 @@ Long documents are split into chunks. Each chunk is scored, and the results are
277
  **Scores.** Each 0-5 score is a decimal such as 2.73: the model estimates how likely each level from 0 to 5 is, and
278
  the score is the average level weighted by those chances. `code_quality` applies when `content_type` is
279
  text_with_code or code_only, or `format` is code_file. `math_quality` applies when `content_type` is math_heavy or
280
- `topic` is math. Otherwise they are `null`. What each level means is in `source1.json` and
281
- [EVALUATION.md Appendix A](EVALUATION.md#appendix-a-rubric-anchors), with the extra rules the teacher was given.
 
 
282
 
283
  **Formula.**
284
 
@@ -299,10 +322,17 @@ fixed rule on the teacher's labels. Stricter lines and what they cost are in
299
  [EVALUATION.md](EVALUATION.md#the-drop-line). You do not have to use `keep`: ranking by `overall`, or your own rules
300
  on single fields, may work better for you.
301
 
 
 
 
 
 
302
  **Long documents.** A document longer than about 7,800 tokens is split into balanced chunks at natural breaks. Each
303
  chunk is scored with a one-line header, as in training (source type, title if given, "Part i of n"). The document gets
304
  the label that covers the most tokens and the token-weighted average of each score. `toxicity` takes the maximum, and
305
- the code and math scores average only the chunks where they apply. `overall` and `keep` are then recomputed.
 
 
306
 
307
  </details>
308
 
@@ -398,8 +428,10 @@ Data sources, filtering, recipe and development details:
398
 
399
  ## Contact
400
 
401
- Questions, corrections and removal requests: the [Community tab](https://huggingface.co/msmth/Source-1/discussions).
402
- Send the URL or dataset of your content and we will remove it from the next release's training data.
 
 
403
 
404
  ## Citation
405
 
 
146
  ## Highlights
147
 
148
  The numbers are rank agreement with an independent proprietary LLM grader: how closely a model puts texts in the same
149
+ order as the grader does (1.0 means the same order, 0 means no link). The grader's grades were never trained on, and on
150
+ the main test set it was given the same instructions as Source-1's teacher.
151
+
152
+ - **Beats each of the 16 public quality scorers we tested, including FineWeb-Edu and propella-1, on every test set it
153
+ was run on** (the 11 English-only scorers were not run on the 12-language exam). On the main test set (495 held-out
154
+ texts in 53 languages): 0.90 vs 0.76 for the best of them, propella-1 4B, a model with 13x more parameters.
155
+ - **Nearly matches its teacher**, the open-weight 27B LLM that labeled its training data, with 1/88 of its parameters:
156
+ 0.90 vs 0.91 on the main test set.
157
  - **Also ahead when judged on educational value alone**, the thing most public scorers were built for.
158
 
159
  *Caveat: the grader scored with Source-1's own rubric (its scoring guide), so this test plays to Source-1's
160
+ strengths. Even a scorer that matched the grader's educational-value scores exactly would reach only 0.87 here. See
161
+ [Limitations](#limitations) and [EVALUATION.md](EVALUATION.md).*
162
 
163
  ## Quick start
164
 
 
181
  print(doc["overall"], doc["keep"]) # a 0-5 score and the keep/drop decision
182
  ```
183
 
184
+ ### For AI agents and scripts
185
+
186
+ - **Load it only through `source1.py`.** `pipeline("text-classification", ...)` and `AutoModel...` classes load the
187
+ backbone without Source-1's 13 trained heads and return meaningless `LABEL_0` / `LABEL_1` scores.
188
+ - **Check the setup** with the `examples/sample.jsonl` command above: its output must equal
189
+ `examples/expected_output.jsonl`.
190
+ - **Output:** one JSON object per document (the 13 fields, `overall`, `keep`, `drop_reasons`). **Exit codes:** 0 done;
191
+ 2 bad arguments or an unreadable input, before the model loads; 1 a bad record during a run.
192
+ - No `trust_remote_code` and no prompts. Loading from a local folder makes no network calls.
193
+
194
  <details>
195
  <summary>More usage: many texts, precision, options, speed, files</summary>
196
 
 
219
  ```
220
 
221
  - **Output.** One flat dict per document with the 13 fields, `overall`, `keep` and `drop_reasons`. It also gives the
222
+ length in tokens, the number of chunks (`parts`), whether a chunk was cut, and per-chunk results (`chunks`). Empty
223
+ text (or only spaces, zero-width or control characters) gives `keep` false, `drop_reasons` `["empty text"]` and
224
+ `null` for `overall` and every field, so leave those records out before sorting by `overall`.
225
  - **Precision.** By default it loads `model.safetensors` (bfloat16, 0.6 GB). For the full float32 weights (1.2 GB),
226
  remove `--exclude` from the download and pass `precision="fp32"`. It computes in bfloat16 on NVIDIA Ampere or newer
227
  GPUs and in float32 elsewhere. float16 is not supported. In bfloat16, scores can shift by up to about 0.04
228
  depending on which texts share a batch; pass `dtype="fp32"` or `batch_tokens=1` to avoid that.
229
  - **Options.** `drop_line` sets your own drop rule, such as `"toxicity >= 4 or spam_seo >= 3 or boilerplate >= 4.5"`.
230
  `max_chunks` scores only some chunks of very long documents, for speed. `revision` pins a Hub version. The command
231
+ line reads JSONL, JSON or plain text (also gzip, bzip2 or xz compressed), and `python source1.py --help` lists
232
+ every flag.
233
  - **Examples and speed.** `examples/` holds six documents and their expected CPU output. A GPU can differ by a few
234
  hundredths. Source-1 scores about 30 chunks (50,000 tokens) per second on one RTX 3090 in bfloat16.
235
+ - **On a CPU.** By default it scores 16,384 tokens per batch there and uses about 3 GB of RAM. On 4 threads it reads
236
+ about 900 tokens per second on typical chunks and about 600 on full-length ones (about 12 seconds per
237
+ 7,000-token chunk). Use `--max-chunks` for long documents.
238
 
239
  **Files**
240
 
 
248
  | `calibration.json` | the drop line and the quality-score offsets, with how they were chosen |
249
  | `tokenizer.json`, `tokenizer_config.json` | mmBERT's tokenizer (same vocabulary and merges, re-saved) |
250
  | `source1.py` | standalone loader, Python API and command line |
251
+ | `requirements.txt` | `torch`, `transformers`, `safetensors`, `tokenizers`, `huggingface_hub` |
252
  | `examples/` | six sample inputs (`sample.jsonl`) and their expected command-line output (`expected_output.jsonl`) |
253
  | `EVALUATION.md` | the full evaluation, the rubric anchors and the training details |
254
  | `images/` | the benchmark chart above |
255
  | `LICENSE`, `NOTICE`, `AUTHORS` | license text, third-party notices and credits, authors |
256
+ | `CREDITS_BOOKS.tsv` | per-work credits for the open-books part of the training data, for books that are not public domain or CC0 (part of `NOTICE`) |
257
 
258
  </details>
259
 
 
279
  scores are `null` when they do not apply. You also get:
280
 
281
  - `overall`: one 0-5 score, the quality scores minus penalties for red flags.
282
+ - `keep`: false (drop) if `toxicity >= 4` or `spam_seo >= 3.5` or `boilerplate >= 4.5`. It checks only these red
283
+ flags, so also set a threshold on `overall`.
284
 
285
+ Long documents are split into chunks. Each chunk is scored, the results are combined into one, and `keep` is decided
286
+ on the combined scores (each chunk's own result is in `chunks`).
287
 
288
  <details>
289
  <summary>All fields in detail, the scoring formula, the drop line and long documents</summary>
 
298
  **Scores.** Each 0-5 score is a decimal such as 2.73: the model estimates how likely each level from 0 to 5 is, and
299
  the score is the average level weighted by those chances. `code_quality` applies when `content_type` is
300
  text_with_code or code_only, or `format` is code_file. `math_quality` applies when `content_type` is math_heavy or
301
+ `topic` is math. Otherwise they are `null`. For a split document they average the chunks where they apply, so they can
302
+ be set even when the document's `content_type` is plain_text. What each level means is in `source1.json` and
303
+ [EVALUATION.md Appendix A](EVALUATION.md#appendix-a-rubric-anchors), with the extra rules that the teacher and the
304
+ held-out grader were given.
305
 
306
  **Formula.**
307
 
 
322
  [EVALUATION.md](EVALUATION.md#the-drop-line). You do not have to use `keep`: ranking by `overall`, or your own rules
323
  on single fields, may work better for you.
324
 
325
+ `keep` applies only this red-flag line, so very short or degenerate text (a single word, an emoji, one letter
326
+ repeated) can still be kept. Combine it with a threshold on `overall`. A drop line can also use `overall`, `tokens`
327
+ and `parts` (the last two for whole documents only), for example
328
+ `"toxicity >= 4 or spam_seo >= 3.5 or boilerplate >= 4.5 or overall < 1"`.
329
+
330
  **Long documents.** A document longer than about 7,800 tokens is split into balanced chunks at natural breaks. Each
331
  chunk is scored with a one-line header, as in training (source type, title if given, "Part i of n"). The document gets
332
  the label that covers the most tokens and the token-weighted average of each score. `toxicity` takes the maximum, and
333
+ the code and math scores average only the chunks where they apply. `overall` and `keep` are then recomputed from these
334
+ combined scores, so a document can be kept even when some of its chunks would be dropped. Each chunk's own scores and
335
+ `keep` are in `chunks`.
336
 
337
  </details>
338
 
 
428
 
429
  ## Contact
430
 
431
+ Questions, corrections and removal requests: open a discussion in the
432
+ [Community tab](https://huggingface.co/msmth/Source-1/discussions). If your request involves personal information, open
433
+ a discussion without the details and we will arrange a private way to reach us. We review every request and, where the
434
+ content is in our training data, exclude it from future versions; published weights cannot be changed.
435
 
436
  ## Citation
437
 
requirements.txt CHANGED
@@ -1,8 +1,11 @@
1
  # source1.py needs Python >= 3.10 and these packages. The minimum versions are the oldest this release was tested
2
  # with (torch 2.11 and 2.14, transformers 5.17), with both weight files: model.safetensors (bfloat16, the default)
3
- # and model.fp32.safetensors (float32). Older versions may work but were not tested. No GPU is needed: on a CPU, or a
4
- # GPU without native bfloat16, the bfloat16 weights are upcast to float32 when loaded.
5
- torch>=2.11
6
- transformers>=5.17
 
 
7
  safetensors>=0.8
8
  tokenizers>=0.23
 
 
1
  # source1.py needs Python >= 3.10 and these packages. The minimum versions are the oldest this release was tested
2
  # with (torch 2.11 and 2.14, transformers 5.17), with both weight files: model.safetensors (bfloat16, the default)
3
+ # and model.fp32.safetensors (float32). Older versions may work but were not tested, and the next major versions
4
+ # are left out until they are. huggingface_hub downloads the model when it is loaded by its Hugging Face repo id.
5
+ # No GPU is needed: on a CPU, or a GPU without native bfloat16, the bfloat16 weights are upcast to float32 when
6
+ # loaded.
7
+ torch>=2.11,<3
8
+ transformers>=5.17,<6
9
  safetensors>=0.8
10
  tokenizers>=0.23
11
+ huggingface_hub>=1.0
source1.py CHANGED
@@ -17,6 +17,7 @@ CPU.
17
 
18
  python source1.py --input docs.jsonl --output scores.jsonl # one JSON object per line, "text" field
19
  python source1.py --input page.txt # a .txt file is one document
 
20
 
21
  Output: one flat dict per document
22
  ----------------------------------
@@ -42,8 +43,8 @@ calibration offsets of calibration.json to the quality scores (about 0.02 at mos
42
 
43
  How a document is scored (the same steps the model was trained and evaluated with)
44
  ------------------------------------------------------------------------------------
45
- 1. ``clean_text``: line endings to "\\n", control characters dropped, Unicode NFC, trailing spaces dropped, at most
46
- two blank lines in a row.
47
  2. ``split_text``: a document longer than 7,808 tokens is split into balanced chunks, each ending at the most
48
  natural boundary near its ideal end (headings, then paragraphs, lines, sentences, spaces; definitions in code).
49
  3. ``build_input``: each chunk gets a one-line header, a blank line, then the chunk text:
@@ -51,7 +52,7 @@ How a document is scored (the same steps the model was trained and evaluated wit
51
  Every training input had a Source line, almost always "dataset record", so that is the default; code files had
52
  "<language> source file" (pass ``code_language="Python"``). Title is added when given, Part when the document has
53
  more than one chunk. ``url`` is accepted but not shown to the model unless ``show_url=True``: no training input had
54
- one. On 1,261 held-out benchmark chunks (495 graded held-out chunks from the test split from the test split and 766 exam chunks),
55
  dropping the whole header moved overall by 0.03 on average (at most 0.655), dropping a title by 0.05 on the chunks
56
  that had one; adding a URL moved it by up to 0.7 and did not improve its rank agreement with the independent
57
  graders of the model card's evaluation.
@@ -85,13 +86,17 @@ from __future__ import annotations
85
 
86
  import argparse
87
  import ast
 
 
88
  import json
 
89
  import math
90
  import operator
91
  import re
92
  import sys
93
  import time
94
  import unicodedata
 
95
  from bisect import bisect_left
96
  from collections.abc import Iterable, Iterator, Mapping
97
  from pathlib import Path
@@ -105,7 +110,8 @@ __version__ = "1.0.0"
105
  MAX_LENGTH = 8192 # tokens per model input, <bos> and <eos> included
106
  CHUNK_TOKENS = 7808 # document tokens per chunk: 8,192 minus 384 kept for the header (as in training)
107
  DEFAULT_SOURCE = "dataset record"
108
- DEFAULT_BATCH_TOKENS = 65536 # padded tokens per forward pass
 
109
  LEVELS = (0, 1, 2, 3, 4, 5)
110
  # The backbone weights of each precision: bfloat16 (the default) and the full-precision float32 copy. The fp32 file
111
  # follows transformers' variant naming (model.<variant>.safetensors), so AutoModel loads it with variant="fp32".
@@ -126,7 +132,9 @@ def hub_files(precision: str = DEFAULT_PRECISION) -> tuple[str, ...]:
126
 
127
 
128
  _JUNK = re.compile("[\x00-\x08\x0b\x0e-\x1f\x7f\ud800-\udfff\ufeff\u200b\ufffe\uffff]")
129
- _TRAILING_WS = re.compile(r"[ \t]+\n")
 
 
130
  _BLANK_LINES = re.compile(r"\n{4,}")
131
  _SPACE_RUN = re.compile(r"[ \t]{2,}")
132
  _BLANK_RUN = re.compile(r"\n(?:[ \t]*\n){3,}")
@@ -134,8 +142,9 @@ _BLANK_RUN = re.compile(r"\n(?:[ \t]*\n){3,}")
134
 
135
  def clean_text(text: str) -> str:
136
  """The document as the scorer sees it before chunking: "\\n" line endings (a form feed counts as a paragraph
137
- break), control characters, zero-width spaces and byte-order marks dropped, Unicode NFC, no trailing spaces,
138
- at most two blank lines in a row, no blank lines at either end."""
 
139
  text = text.replace("\r\n", "\n").replace("\r", "\n").replace("\x0c", "\n\n")
140
  text = _JUNK.sub("", text)
141
  text = unicodedata.normalize("NFC", text)
@@ -296,25 +305,107 @@ _CMP = {ast.Eq: operator.eq, ast.NotEq: operator.ne, ast.Lt: operator.lt, ast.Lt
296
  ast.Gt: operator.gt, ast.GtE: operator.ge}
297
 
298
 
 
 
 
 
299
  class DropLine:
300
- """A drop line such as ``toxicity >= 4 or spam_seo >= 3.5 or boilerplate >= 4.5``: field names, numbers,
301
- comparisons, ``and`` / ``or`` / ``not`` and parentheses. A comparison with a missing score (None) is false.
302
- A chunk or document is kept when the line does not match."""
 
 
 
 
 
 
 
303
 
304
  _NODES = (ast.Expression, ast.BoolOp, ast.And, ast.Or, ast.UnaryOp, ast.Not, ast.USub, ast.Compare, ast.Name,
305
  ast.Load, ast.Constant, *_CMP)
306
 
307
- def __init__(self, source: str, names: Iterable[str]):
308
  self.source = source.strip()
309
- tree = ast.parse(self.source, mode="eval")
 
 
 
 
 
310
  for node in ast.walk(tree):
311
  if not isinstance(node, self._NODES):
312
- raise ValueError(f"{type(node).__name__} is not allowed in a drop line: {source!r}")
313
- if isinstance(node, ast.Name) and node.id not in set(names):
314
- raise ValueError(f"unknown name {node.id!r} in drop line {source!r}")
315
  self.tree = tree.body
 
 
 
 
316
  top_or = isinstance(self.tree, ast.BoolOp) and isinstance(self.tree.op, ast.Or)
317
  self.terms = list(self.tree.values) if top_or else [self.tree]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
318
 
319
  def reasons(self, env: dict) -> list[str]:
320
  """The terms of the line that match ``env`` (each ``or`` branch on its own); [] means keep."""
@@ -341,16 +432,11 @@ class DropLine:
341
  left = self._eval(node.left, env)
342
  for op, comp in zip(node.ops, node.comparators):
343
  right = self._eval(comp, env)
344
- if not isinstance(op, (ast.Eq, ast.NotEq)) and (left is None or right is None):
345
- return False
346
- try:
347
- if not _CMP[type(op)](left, right):
348
- return False
349
- except TypeError:
350
- return False
351
  left = right
352
  return True
353
- raise ValueError(f"unsupported drop-line node {type(node).__name__}")
354
 
355
 
356
  def _r3(x: float | None) -> float | None:
@@ -567,6 +653,23 @@ def _load_calibration(path: Path, drop_line: str | None, apply_offsets: bool) ->
567
  return calibration
568
 
569
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
570
  class Source1(nn.Module):
571
  """Source-1: the mmBERT-base encoder, mean pooling and one linear head per field. Build it with
572
  ``Source1.from_pretrained``; score documents with ``score`` / ``score_batch``."""
@@ -590,15 +693,7 @@ class Source1(nn.Module):
590
  self.labels = {name: list(spec["values"]) for name, spec in self.schema["labels"].items()}
591
  self.fields = list(self.labels) + [n for g in ("quality", "red_flags", "gated") for n in self.schema.get(g, {})]
592
  self.calibration = calibration or {}
593
- default_line = " or ".join((self.schema.get("composite") or {}).get("hard_filters") or []) or "False"
594
- if drop_line in (None, "calibrated"):
595
- drop_line = (self.calibration.get("drop_line") or {}).get("line")
596
- if not drop_line:
597
- raise ValueError("no calibrated drop line (calibration.json missing or incomplete); pass "
598
- "drop_line='default' or your own drop line")
599
- elif drop_line == "default":
600
- drop_line = default_line
601
- self.drop_line = DropLine(drop_line, set(self.fields) | {"overall", "tokens", "parts"})
602
  self.offsets = {k: float(v["offset"]) for k, v in (self.calibration.get("offsets") or {}).items()}
603
  self.apply_offsets = apply_offsets
604
  self.show_url = show_url
@@ -658,11 +753,21 @@ class Source1(nn.Module):
658
  raise FileNotFoundError(f"{path} lacks {', '.join(missing)}")
659
  config = json.loads((path / "source1.json").read_text(encoding="utf-8"))
660
  calibration = _load_calibration(path, drop_line, apply_offsets)
 
661
  tok = Tokenizer.from_file(str(path / "tokenizer.json"))
662
  tok.no_truncation()
663
  tok.no_padding()
664
  variant = {"variant": "fp32"} if precision == "fp32" else {}
665
- backbone = AutoModel.from_pretrained(str(path), **{_dtype_kwarg(): dtype}, **variant, **backbone_kwargs)
 
 
 
 
 
 
 
 
 
666
  model = cls(backbone, tok, config, calibration, drop_line, apply_offsets, show_url)
667
  model.heads.load_state_dict(load_file(str(path / "heads.safetensors")))
668
  model.heads.to(dtype)
@@ -754,12 +859,19 @@ class Source1(nn.Module):
754
  rec["drop_reasons"] = reasons
755
  return rec
756
 
757
- def score_inputs(self, inputs: list[str], batch_tokens: int = DEFAULT_BATCH_TOKENS) -> list[dict]:
 
 
 
 
758
  """Score ready-made model inputs (``build_input`` output: header, blank line, chunk text), one dict per
759
  input with the 13 fields, overall, keep, drop_reasons, input_tokens and truncated.
760
 
761
- Inputs are sorted by length and batched with at most ``batch_tokens`` padded tokens per forward pass;
762
- a batch that runs out of GPU memory is split in half and retried."""
 
 
 
763
  encoded = self.encode(inputs)
764
  ids = [e[0] for e in encoded]
765
  results: list[dict | None] = [None] * len(ids)
@@ -789,13 +901,13 @@ class Source1(nn.Module):
789
  run(batch)
790
  return results # type: ignore[return-value]
791
 
792
- def score_batch(self, docs: Iterable[str | bytes | dict], batch_tokens: int = DEFAULT_BATCH_TOKENS, *,
793
  text_field: str = "text", max_chunks: int = 0) -> list[dict]:
794
  """Score a list of documents: strings (bytes are decoded as UTF-8), or dicts with the text under
795
  ``text_field`` and optionally ``title``, ``url``, ``source_type`` and ``code_language`` (see ``score``).
796
  A None text is scored as an empty document (keep False, drop_reasons ["empty text"]). Chunks of all
797
  documents are batched together. ``max_chunks`` > 0 scores only that many evenly spaced chunks of a long
798
- document (the Part numbers still count every chunk).
799
 
800
  In bfloat16 a document's scores can shift slightly (up to about 0.04 on overall, 0.10 on a single field)
801
  depending on which other documents share its batch, because the batch shape changes the kernels' rounding;
@@ -831,7 +943,7 @@ class Source1(nn.Module):
831
 
832
  def score(self, text: str | bytes, title: str | None = None, url: str | None = None, *,
833
  source_type: str | None = None, code_language: str | None = None, max_chunks: int = 0,
834
- batch_tokens: int = DEFAULT_BATCH_TOKENS) -> dict:
835
  """Score one document (a str; bytes are decoded as UTF-8; anything else raises TypeError).
836
 
837
  title: shown to the model in the header when given (as in training, where about a quarter of inputs had one).
@@ -869,14 +981,108 @@ class Source1(nn.Module):
869
 
870
 
871
  class BadInput(ValueError):
872
- """A record of the input file that cannot be scored (the message starts with file:line)."""
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
873
 
874
 
875
  def _read_docs(path: Path, text_field: str, skip_bad: bool = False) -> Iterator[dict]:
876
  """Documents of an input file: .jsonl / .ndjson (one JSON object per line), .json (a JSON array of objects, one
877
- object, or JSON Lines), or any other file as one plain-text document. Invalid UTF-8 is replaced (with a warning
878
- for JSON). A bad record raises BadInput, or with ``skip_bad`` is reported on stderr and skipped. A null text is
879
- kept and scored as an empty document."""
 
 
880
 
881
  def usable(rec: Any, where: str) -> bool:
882
  if not isinstance(rec, dict):
@@ -885,6 +1091,8 @@ def _read_docs(path: Path, text_field: str, skip_bad: bool = False) -> Iterator[
885
  problem = f"no {text_field!r} field"
886
  elif rec[text_field] is not None and not isinstance(rec[text_field], str):
887
  problem = f"{text_field!r} is a {type(rec[text_field]).__name__}, not a string"
 
 
888
  else:
889
  return True
890
  if not skip_bad:
@@ -900,37 +1108,52 @@ def _read_docs(path: Path, text_field: str, skip_bad: bool = False) -> Iterator[
900
  file=sys.stderr)
901
  return raw.decode("utf-8", errors="replace")
902
 
903
- suffix = path.suffix.lower()
 
 
 
 
 
 
 
 
 
 
904
  if suffix not in (".jsonl", ".ndjson", ".json"):
905
- yield {text_field: path.read_text(encoding="utf-8", errors="replace"), "id": path.name}
906
  return
907
  if suffix == ".json":
908
  try:
909
- data = json.loads(decode(path.read_bytes(), str(path)).lstrip(""))
910
- except json.JSONDecodeError:
911
- data = None # not a single JSON value: read the file as JSON Lines below
912
  if data is not None:
913
  for n, rec in enumerate(data if isinstance(data, list) else [data]):
914
  if usable(rec, f"{path}[{n}]"):
915
  yield rec
916
  return
917
- with open(path, "rb") as f:
918
- for n, raw in enumerate(f, 1):
919
- where = f"{path}:{n}"
920
- line = decode(raw, where)
921
- if n == 1:
922
- line = line.lstrip("")
923
- if not line.strip():
924
- continue
925
- try:
926
- rec = json.loads(line)
927
- except json.JSONDecodeError as e:
928
- if not skip_bad:
929
- raise BadInput(f"{where}: invalid JSON ({e.msg} at column {e.colno})") from None
930
- print(f"source1: skipped {where}: invalid JSON ({e.msg} at column {e.colno})", file=sys.stderr)
931
- continue
932
- if usable(rec, where):
933
- yield rec
 
 
 
 
 
934
 
935
 
936
  def main(argv: list[str] | None = None) -> int:
@@ -938,10 +1161,11 @@ def main(argv: list[str] | None = None) -> int:
938
  p.add_argument("--model", default=str(Path(__file__).resolve().parent),
939
  help="Source-1 directory or Hugging Face repo id (default: this file's directory)")
940
  p.add_argument("--input", required=True, help=".jsonl (one document per line), .json (an array of objects) or "
941
- "a text file (one document)")
942
  p.add_argument("--text-field", default="text", help="JSON field holding the text (default: text); "
943
  "title, url, source_type and code_language fields are used when present")
944
- p.add_argument("--output", help="output .jsonl (default: standard output)")
 
945
  p.add_argument("--skip-bad", action="store_true", help="skip (and report on stderr) records that are not valid "
946
  "JSON objects with a string text, instead of stopping")
947
  p.add_argument("--revision", help="branch, tag or commit, when --model is a Hugging Face repo id")
@@ -951,7 +1175,8 @@ def main(argv: list[str] | None = None) -> int:
951
  p.add_argument("--dtype", choices=DTYPE_CHOICES, default="auto", help="what to compute in; auto (default): "
952
  "bfloat16 on a GPU with native bfloat16, else float32 (bf16 weights upcast); float16 is not "
953
  "supported")
954
- p.add_argument("--batch-tokens", type=int, default=DEFAULT_BATCH_TOKENS, help="padded tokens per forward pass")
 
955
  p.add_argument("--max-chunks", type=int, default=0, help="score at most N evenly spaced chunks per document")
956
  p.add_argument("--drop-line", help='"calibrated" (default), "default" (the schema\'s hard filters) or an '
957
  "expression such as 'toxicity >= 4 or spam_seo >= 3'")
@@ -960,50 +1185,109 @@ def main(argv: list[str] | None = None) -> int:
960
  p.add_argument("--no-chunks", action="store_true", help="leave out the per-chunk list of split documents")
961
  p.add_argument("--group", type=int, default=256, help="documents scored together")
962
  args = p.parse_args(argv)
963
- if not Path(args.input).is_file():
 
964
  p.error(f"--input {args.input}: no such file")
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
965
 
966
  t0 = time.time()
967
- model = Source1.from_pretrained(args.model, device=args.device, dtype=args.dtype, precision=args.precision,
968
- drop_line=args.drop_line, apply_offsets=args.apply_offsets,
969
- show_url=args.show_url, revision=args.revision)
 
 
 
970
  compute = str(next(model.parameters()).dtype).replace("torch.", "")
971
  print(f"source1: loaded {model.weights_file} ({model.weights_dtype or '?'} weights) on {model.device}, computing "
972
  f"in {compute}, in {time.time() - t0:.1f} s; drop line: {model.drop_line.source}", file=sys.stderr)
973
- out = open(args.output, "w", encoding="utf-8") if args.output else sys.stdout
974
- done = 0
975
  t0 = time.time()
976
 
977
  def flush(group: list[dict]) -> None:
978
- nonlocal done
979
  for rec, res in zip(group, model.score_batch(group, args.batch_tokens, text_field=args.text_field,
980
  max_chunks=args.max_chunks)):
 
981
  if args.no_chunks:
982
  res.pop("chunks", None)
983
  if "id" in rec:
984
  res = {"id": rec["id"], **res}
985
- out.write(json.dumps(res, ensure_ascii=False, allow_nan=False) + "\n")
986
- done += len(group)
 
 
 
 
 
 
 
 
 
987
  print(f"source1: {done:,} documents, {done / max(time.time() - t0, 1e-9):.1f}/s", file=sys.stderr)
988
 
 
 
 
 
 
 
 
 
 
 
 
 
989
  try:
990
- group: list[dict] = []
991
  try:
992
- for rec in _read_docs(Path(args.input), args.text_field, args.skip_bad):
 
 
 
 
 
 
 
 
993
  group.append(rec)
994
  if len(group) >= args.group:
995
  flush(group)
996
  group = []
997
- except BadInput as e:
998
  if group:
999
- flush(group) # the documents read before the bad record are still scored and written
1000
- raise SystemExit(f"source1: {e}. The {done:,} documents before it were written; --skip-bad skips "
1001
- "bad records") from None
1002
- if group:
1003
- flush(group)
1004
- finally:
1005
- if out is not sys.stdout:
1006
- out.close()
 
 
 
 
 
1007
  return 0
1008
 
1009
 
 
17
 
18
  python source1.py --input docs.jsonl --output scores.jsonl # one JSON object per line, "text" field
19
  python source1.py --input page.txt # a .txt file is one document
20
+ python source1.py --input docs.jsonl.gz # gzip, bzip2 and xz files are decompressed
21
 
22
  Output: one flat dict per document
23
  ----------------------------------
 
43
 
44
  How a document is scored (the same steps the model was trained and evaluated with)
45
  ------------------------------------------------------------------------------------
46
+ 1. ``clean_text``: line endings to "\\n", ASCII control characters other than tab and newline dropped, Unicode NFC,
47
+ trailing spaces dropped, at most two blank lines in a row.
48
  2. ``split_text``: a document longer than 7,808 tokens is split into balanced chunks, each ending at the most
49
  natural boundary near its ideal end (headings, then paragraphs, lines, sentences, spaces; definitions in code).
50
  3. ``build_input``: each chunk gets a one-line header, a blank line, then the chunk text:
 
52
  Every training input had a Source line, almost always "dataset record", so that is the default; code files had
53
  "<language> source file" (pass ``code_language="Python"``). Title is added when given, Part when the document has
54
  more than one chunk. ``url`` is accepted but not shown to the model unless ``show_url=True``: no training input had
55
+ one. On 1,261 held-out benchmark chunks (495 graded held-out chunks from the test split and 766 exam chunks),
56
  dropping the whole header moved overall by 0.03 on average (at most 0.655), dropping a title by 0.05 on the chunks
57
  that had one; adding a URL moved it by up to 0.7 and did not improve its rank agreement with the independent
58
  graders of the model card's evaluation.
 
86
 
87
  import argparse
88
  import ast
89
+ import bz2
90
+ import gzip
91
  import json
92
+ import lzma
93
  import math
94
  import operator
95
  import re
96
  import sys
97
  import time
98
  import unicodedata
99
+ import zlib
100
  from bisect import bisect_left
101
  from collections.abc import Iterable, Iterator, Mapping
102
  from pathlib import Path
 
110
  MAX_LENGTH = 8192 # tokens per model input, <bos> and <eos> included
111
  CHUNK_TOKENS = 7808 # document tokens per chunk: 8,192 minus 384 kept for the header (as in training)
112
  DEFAULT_SOURCE = "dataset record"
113
+ DEFAULT_BATCH_TOKENS = 65536 # padded tokens per forward pass on a GPU
114
+ DEFAULT_BATCH_TOKENS_CPU = 16384 # on a CPU: a quarter of the memory, and no slower there
115
  LEVELS = (0, 1, 2, 3, 4, 5)
116
  # The backbone weights of each precision: bfloat16 (the default) and the full-precision float32 copy. The fp32 file
117
  # follows transformers' variant naming (model.<variant>.safetensors), so AutoModel loads it with variant="fp32".
 
132
 
133
 
134
  _JUNK = re.compile("[\x00-\x08\x0b\x0e-\x1f\x7f\ud800-\udfff\ufeff\u200b\ufffe\uffff]")
135
+ # The lookbehind keeps this linear: without it, a long run of spaces that ends in no newline is scanned again from
136
+ # each of its positions, so a page of spaces takes quadratic time. The result is the same.
137
+ _TRAILING_WS = re.compile(r"(?<![ \t])[ \t]+\n")
138
  _BLANK_LINES = re.compile(r"\n{4,}")
139
  _SPACE_RUN = re.compile(r"[ \t]{2,}")
140
  _BLANK_RUN = re.compile(r"\n(?:[ \t]*\n){3,}")
 
142
 
143
  def clean_text(text: str) -> str:
144
  """The document as the scorer sees it before chunking: "\\n" line endings (a form feed counts as a paragraph
145
+ break), ASCII control characters other than tab and newline, lone surrogates, zero-width spaces and byte-order
146
+ marks dropped (C1 control characters such as U+0085 are kept), Unicode NFC, no trailing spaces, at most two
147
+ blank lines in a row, no blank lines at either end."""
148
  text = text.replace("\r\n", "\n").replace("\r", "\n").replace("\x0c", "\n\n")
149
  text = _JUNK.sub("", text)
150
  text = unicodedata.normalize("NFC", text)
 
305
  ast.Gt: operator.gt, ast.GtE: operator.ge}
306
 
307
 
308
+ class DropLineError(ValueError):
309
+ """A drop line that is not valid (see ``DropLine``)."""
310
+
311
+
312
  class DropLine:
313
+ """A drop line such as ``toxicity >= 4 or spam_seo >= 3.5 or boilerplate >= 4.5``: comparisons of a score (or
314
+ ``overall``, ``tokens``, ``parts``) with a number or another score, and of a label with one of its values
315
+ (``format == 'news'``; ``==`` and ``!=`` only), joined with ``and`` / ``or`` / ``not`` and parentheses; ``True``
316
+ and ``False`` match always and never. A comparison with a missing score (None) is false, ``!=`` included. A
317
+ chunk or document is kept when the line does not match.
318
+
319
+ ``names`` are the fields the line may use and ``labels`` maps each label field to its values; the other names
320
+ are numbers. The line is checked when it is created: a bare field name or constant used as a condition, a string
321
+ that is not a value of its label, a score compared with a string or a label with a number raise DropLineError
322
+ (a ValueError)."""
323
 
324
  _NODES = (ast.Expression, ast.BoolOp, ast.And, ast.Or, ast.UnaryOp, ast.Not, ast.USub, ast.Compare, ast.Name,
325
  ast.Load, ast.Constant, *_CMP)
326
 
327
+ def __init__(self, source: str, names: Iterable[str], labels: Mapping[str, Iterable[str]] | None = None):
328
  self.source = source.strip()
329
+ self.names = set(names)
330
+ self.labels = {name: list(values) for name, values in (labels or {}).items()}
331
+ try:
332
+ tree = ast.parse(self.source, mode="eval")
333
+ except (SyntaxError, ValueError, RecursionError, MemoryError) as e:
334
+ raise DropLineError(f"drop line {source!r} is not a valid expression ({e})") from None
335
  for node in ast.walk(tree):
336
  if not isinstance(node, self._NODES):
337
+ raise DropLineError(f"{type(node).__name__} is not allowed in a drop line: {source!r}")
338
+ if isinstance(node, ast.Name) and node.id not in self.names:
339
+ raise DropLineError(f"unknown name {node.id!r} in drop line {source!r}")
340
  self.tree = tree.body
341
+ try:
342
+ self._condition(self.tree)
343
+ except RecursionError:
344
+ raise DropLineError(f"drop line {source!r} is nested too deeply") from None
345
  top_or = isinstance(self.tree, ast.BoolOp) and isinstance(self.tree.op, ast.Or)
346
  self.terms = list(self.tree.values) if top_or else [self.tree]
347
+ # Evaluated once on stand-in scores, and once with every score missing, so that a line that cannot be
348
+ # evaluated fails here and not in the middle of a run.
349
+ try:
350
+ self.reasons({n: self.labels[n][0] if self.labels.get(n) else 0.0 for n in self.names})
351
+ self.reasons({})
352
+ except RecursionError:
353
+ raise DropLineError(f"drop line {source!r} is nested too deeply") from None
354
+
355
+ def _fail(self, node: ast.AST, problem: str) -> None:
356
+ raise DropLineError(f"{ast.unparse(node)!r} {problem}, in drop line {self.source!r}")
357
+
358
+ def _condition(self, node: ast.AST) -> None:
359
+ """Check that ``node`` is a condition: a comparison, True / False, or and / or / not of conditions."""
360
+ if isinstance(node, ast.BoolOp):
361
+ for value in node.values:
362
+ self._condition(value)
363
+ elif isinstance(node, ast.UnaryOp) and isinstance(node.op, ast.Not):
364
+ self._condition(node.operand)
365
+ elif isinstance(node, ast.Compare):
366
+ kinds = [self._operand(x) for x in (node.left, *node.comparators)]
367
+ for op, (a, b) in zip(node.ops, zip(kinds, kinds[1:])):
368
+ self._check_pair(node, op, a, b)
369
+ elif not (isinstance(node, ast.Constant) and isinstance(node.value, bool)):
370
+ self._fail(node, "is not a condition: compare it with something, as in 'toxicity >= 4'")
371
+
372
+ def _operand(self, node: ast.AST) -> tuple[str, Any]:
373
+ """What one side of a comparison is: ("label", name), ("str", value), ("num", name) for a numeric field or
374
+ ("num", None) for a number."""
375
+ if isinstance(node, ast.Name):
376
+ return ("label", node.id) if node.id in self.labels else ("num", node.id)
377
+ if isinstance(node, ast.UnaryOp) and isinstance(node.op, ast.USub):
378
+ kind = self._operand(node.operand)
379
+ if kind[0] != "num":
380
+ self._fail(node, "negates something that is not a number")
381
+ return kind
382
+ if isinstance(node, ast.Constant):
383
+ v = node.value
384
+ if isinstance(v, str):
385
+ return ("str", v)
386
+ if (isinstance(v, int) and not isinstance(v, bool)) or (isinstance(v, float) and math.isfinite(v)):
387
+ return ("num", None)
388
+ self._fail(node, "is not allowed in a comparison: use a finite number or a label value in quotes")
389
+ self._fail(node, "is not allowed in a comparison: compare fields, numbers and label values")
390
+ raise AssertionError # not reached
391
+
392
+ def _check_pair(self, node: ast.Compare, op: ast.cmpop, a: tuple[str, Any], b: tuple[str, Any]) -> None:
393
+ kinds = {a[0], b[0]}
394
+ if all(k == "str" or name is None for k, name in (a, b)):
395
+ self._fail(node, "compares two constants")
396
+ if kinds == {"num"}:
397
+ return
398
+ if kinds == {"label", "str"}:
399
+ (_, label), (_, value) = (a, b) if a[0] == "label" else (b, a)
400
+ if not isinstance(op, (ast.Eq, ast.NotEq)):
401
+ self._fail(node, f"compares the label {label} by order: use == or !=")
402
+ if value not in self.labels[label]:
403
+ self._fail(node, f"uses {value!r}, which is not a value of {label} (one of: "
404
+ f"{', '.join(self.labels[label])})")
405
+ return
406
+ if kinds == {"label"}:
407
+ self._fail(node, "compares two labels: compare a label with one of its values, as in format == 'news'")
408
+ self._fail(node, "compares a score with a string, or a label with a number")
409
 
410
  def reasons(self, env: dict) -> list[str]:
411
  """The terms of the line that match ``env`` (each ``or`` branch on its own); [] means keep."""
 
432
  left = self._eval(node.left, env)
433
  for op, comp in zip(node.ops, node.comparators):
434
  right = self._eval(comp, env)
435
+ if left is None or right is None or not _CMP[type(op)](left, right):
436
+ return False # a comparison with a missing value is false, != included
 
 
 
 
 
437
  left = right
438
  return True
439
+ raise DropLineError(f"unsupported drop-line node {type(node).__name__}")
440
 
441
 
442
  def _r3(x: float | None) -> float | None:
 
653
  return calibration
654
 
655
 
656
+ def _make_drop_line(config: dict, calibration: dict | None, drop_line: str | None = None) -> DropLine:
657
+ """The ``DropLine`` of ``drop_line`` for the schema of source1.json (``config``): None or "calibrated" =
658
+ calibration.json's line, "default" = the schema's hard filters, anything else = the expression itself, checked
659
+ against the schema's fields and label values (ValueError when it is not a valid drop line)."""
660
+ schema = config["schema"]
661
+ if drop_line in (None, "calibrated"):
662
+ drop_line = ((calibration or {}).get("drop_line") or {}).get("line")
663
+ if not drop_line:
664
+ raise ValueError("no calibrated drop line (calibration.json missing or incomplete); pass "
665
+ "drop_line='default' or your own drop line")
666
+ elif drop_line == "default":
667
+ drop_line = " or ".join((schema.get("composite") or {}).get("hard_filters") or []) or "False"
668
+ labels = {name: list(spec["values"]) for name, spec in schema["labels"].items()}
669
+ numbers = [n for g in ("quality", "red_flags", "gated") for n in schema.get(g, {})] + ["overall", "tokens", "parts"]
670
+ return DropLine(drop_line, [*labels, *numbers], labels)
671
+
672
+
673
  class Source1(nn.Module):
674
  """Source-1: the mmBERT-base encoder, mean pooling and one linear head per field. Build it with
675
  ``Source1.from_pretrained``; score documents with ``score`` / ``score_batch``."""
 
693
  self.labels = {name: list(spec["values"]) for name, spec in self.schema["labels"].items()}
694
  self.fields = list(self.labels) + [n for g in ("quality", "red_flags", "gated") for n in self.schema.get(g, {})]
695
  self.calibration = calibration or {}
696
+ self.drop_line = _make_drop_line(config, self.calibration, drop_line)
 
 
 
 
 
 
 
 
697
  self.offsets = {k: float(v["offset"]) for k, v in (self.calibration.get("offsets") or {}).items()}
698
  self.apply_offsets = apply_offsets
699
  self.show_url = show_url
 
753
  raise FileNotFoundError(f"{path} lacks {', '.join(missing)}")
754
  config = json.loads((path / "source1.json").read_text(encoding="utf-8"))
755
  calibration = _load_calibration(path, drop_line, apply_offsets)
756
+ _make_drop_line(config, calibration, drop_line) # a bad drop line fails here, before the weights load
757
  tok = Tokenizer.from_file(str(path / "tokenizer.json"))
758
  tok.no_truncation()
759
  tok.no_padding()
760
  variant = {"variant": "fp32"} if precision == "fp32" else {}
761
+ backbone_kwargs.pop("output_loading_info", None) # always requested, to check the weights
762
+ backbone, info = AutoModel.from_pretrained(str(path), **{_dtype_kwarg(): dtype}, **variant, **backbone_kwargs,
763
+ output_loading_info=True)
764
+ keys = ("missing_keys", "unexpected_keys", "mismatched_keys", "error_msgs")
765
+ bad = {k: info[k] for k in keys if info.get(k)}
766
+ if bad: # transformers would otherwise fill the missing weights with random values, without an error
767
+ detail = "; ".join(f"{len(v)} {k.replace('_', ' ')} ({', '.join(sorted(map(str, v))[:3])}"
768
+ f"{', ...' if len(v) > 3 else ''})" for k, v in bad.items())
769
+ raise RuntimeError(f"Source-1: the backbone weights in {path / weights} do not match the model: {detail}. "
770
+ "Download the weights again, and use the package versions in requirements.txt")
771
  model = cls(backbone, tok, config, calibration, drop_line, apply_offsets, show_url)
772
  model.heads.load_state_dict(load_file(str(path / "heads.safetensors")))
773
  model.heads.to(dtype)
 
859
  rec["drop_reasons"] = reasons
860
  return rec
861
 
862
+ def default_batch_tokens(self) -> int:
863
+ """The padded tokens per forward pass when ``batch_tokens`` is None: 16,384 on a CPU, 65,536 on a GPU."""
864
+ return DEFAULT_BATCH_TOKENS_CPU if self.device.type == "cpu" else DEFAULT_BATCH_TOKENS
865
+
866
+ def score_inputs(self, inputs: list[str], batch_tokens: int | None = None) -> list[dict]:
867
  """Score ready-made model inputs (``build_input`` output: header, blank line, chunk text), one dict per
868
  input with the 13 fields, overall, keep, drop_reasons, input_tokens and truncated.
869
 
870
+ Inputs are sorted by length and batched with at most ``batch_tokens`` padded tokens per forward pass
871
+ (default: 16,384 on a CPU, 65,536 on a GPU); a batch that runs out of GPU memory is split in half and
872
+ retried."""
873
+ if batch_tokens is None:
874
+ batch_tokens = self.default_batch_tokens()
875
  encoded = self.encode(inputs)
876
  ids = [e[0] for e in encoded]
877
  results: list[dict | None] = [None] * len(ids)
 
901
  run(batch)
902
  return results # type: ignore[return-value]
903
 
904
+ def score_batch(self, docs: Iterable[str | bytes | dict], batch_tokens: int | None = None, *,
905
  text_field: str = "text", max_chunks: int = 0) -> list[dict]:
906
  """Score a list of documents: strings (bytes are decoded as UTF-8), or dicts with the text under
907
  ``text_field`` and optionally ``title``, ``url``, ``source_type`` and ``code_language`` (see ``score``).
908
  A None text is scored as an empty document (keep False, drop_reasons ["empty text"]). Chunks of all
909
  documents are batched together. ``max_chunks`` > 0 scores only that many evenly spaced chunks of a long
910
+ document (the Part numbers still count every chunk). ``batch_tokens``: see ``score_inputs``.
911
 
912
  In bfloat16 a document's scores can shift slightly (up to about 0.04 on overall, 0.10 on a single field)
913
  depending on which other documents share its batch, because the batch shape changes the kernels' rounding;
 
943
 
944
  def score(self, text: str | bytes, title: str | None = None, url: str | None = None, *,
945
  source_type: str | None = None, code_language: str | None = None, max_chunks: int = 0,
946
+ batch_tokens: int | None = None) -> dict:
947
  """Score one document (a str; bytes are decoded as UTF-8; anything else raises TypeError).
948
 
949
  title: shown to the model in the header when given (as in training, where about a quarter of inputs had one).
 
981
 
982
 
983
  class BadInput(ValueError):
984
+ """A record of the input file that cannot be scored (the message starts with file:line), or, with
985
+ ``skippable=False``, an input file that cannot be read at all."""
986
+
987
+ def __init__(self, message: str, skippable: bool = True):
988
+ super().__init__(message)
989
+ self.skippable = skippable
990
+
991
+
992
+ # Compressed inputs read transparently, recognized by their first bytes whatever their name.
993
+ _DECOMPRESS = {"gzip": gzip.open, "bzip2": bz2.open, "xz": lzma.open}
994
+ _COMPRESSED_SUFFIXES = (".gz", ".bz2", ".xz") # docs.jsonl.gz is read as .jsonl
995
+ _UNREADABLE = {".zst": "is zstd-compressed: decompress it first (zstd -d)",
996
+ ".zstd": "is zstd-compressed: decompress it first (zstd -d)",
997
+ ".parquet": "is a Parquet file: convert it to JSON Lines first",
998
+ ".zip": "is a zip archive: extract it first",
999
+ ".7z": "is a 7z archive: extract it first"}
1000
+ _READ_ERRORS = (OSError, EOFError, zlib.error, lzma.LZMAError) # what damaged or truncated compressed files raise
1001
+ _SNIFF_BYTES = 65536 # how much of an input file _input_problem looks at
1002
+ _MAX_INVALID = 0.2 # share of invalid UTF-8 sequences among the characters above which an input file is refused
1003
+ _MAX_NUL = 0.01 # share of NUL bytes above which an input file is refused as binary
1004
+
1005
+
1006
+ def _compression(head: bytes) -> str | None:
1007
+ """The compression that a file's first bytes show: "gzip", "bzip2", "xz", "zstd" or None."""
1008
+ if head.startswith(b"\x1f\x8b"):
1009
+ return "gzip"
1010
+ if head[:3] == b"BZh" and b"1" <= head[3:4] <= b"9" and head[4:10] in (b"1AY&SY", b"\x17rE8P\x90"):
1011
+ return "bzip2"
1012
+ if head.startswith(b"\xfd7zXZ\x00"):
1013
+ return "xz"
1014
+ if head.startswith(b"\x28\xb5\x2f\xfd"):
1015
+ return "zstd"
1016
+ return None
1017
+
1018
+
1019
+ def _open_input(path: Path) -> Any:
1020
+ """``path`` opened for reading bytes, decompressed when it is a gzip, bzip2 or xz file."""
1021
+ with open(path, "rb") as f:
1022
+ head = f.read(10)
1023
+ return _DECOMPRESS.get(_compression(head) or "", open)(path, "rb")
1024
+
1025
+
1026
+ def _format_suffix(path: Path) -> str:
1027
+ """The suffix that says how to read ``path``: its last one, or the one before .gz, .bz2 or .xz."""
1028
+ suffixes = [s.lower() for s in path.suffixes]
1029
+ if suffixes and suffixes[-1] in _COMPRESSED_SUFFIXES:
1030
+ suffixes.pop()
1031
+ return suffixes[-1] if suffixes else ""
1032
+
1033
+
1034
+ def _input_problem(path: Path) -> str | None:
1035
+ """Why ``path`` cannot be scored as UTF-8 text or JSON (zstd, Parquet, an archive, binary, UTF-16, mostly invalid
1036
+ UTF-8, damaged), judged from its name and its first 64 KiB once decompressed; None when it looks readable."""
1037
+ if path.suffix.lower() in _UNREADABLE:
1038
+ return _UNREADABLE[path.suffix.lower()]
1039
+ try:
1040
+ with _open_input(path) as f:
1041
+ head = f.read(_SNIFF_BYTES)
1042
+ except _READ_ERRORS as e:
1043
+ return f"cannot be read ({e})"
1044
+ if _compression(head) == "zstd":
1045
+ return _UNREADABLE[".zst"]
1046
+ if head.startswith((b"\xff\xfe", b"\xfe\xff")):
1047
+ return "is UTF-16 or UTF-32: save it as UTF-8"
1048
+ if head.count(b"\x00") > _MAX_NUL * len(head): # a stray NUL in a text is dropped like other control characters
1049
+ return "holds NUL bytes, so it is not UTF-8 text (binary, UTF-16, or compressed in a format not read here)"
1050
+ text = head.decode("utf-8", errors="replace")
1051
+ invalid = text.count("\ufffd") - head.count("\ufffd".encode()) # a U+FFFD already in the text is valid UTF-8
1052
+ if invalid >= 8 and invalid > _MAX_INVALID * len(text):
1053
+ return f"is not UTF-8 text: {invalid / len(text):.0%} of its first {len(text):,} characters are invalid UTF-8"
1054
+ return None
1055
+
1056
+
1057
+ def _json_error(e: BaseException) -> str:
1058
+ """Why json.loads failed, in a few words."""
1059
+ if isinstance(e, json.JSONDecodeError):
1060
+ return f"{e.msg} at column {e.colno}"
1061
+ if isinstance(e, RecursionError):
1062
+ return "nested too deeply"
1063
+ return str(e).split(";")[0] # e.g. an integer of more than 4,300 digits
1064
+
1065
+
1066
+ def _unwritable(value: Any) -> str | None:
1067
+ """Why ``value`` cannot be written to the JSON Lines output, or None when it can."""
1068
+ try:
1069
+ json.dumps(value, ensure_ascii=False, allow_nan=False).encode("utf-8")
1070
+ except UnicodeEncodeError:
1071
+ return "holds a lone surrogate, which UTF-8 cannot encode"
1072
+ except ValueError:
1073
+ return "holds NaN or an infinite number, which JSON cannot hold"
1074
+ except RecursionError:
1075
+ return "is nested too deeply"
1076
+ return None
1077
 
1078
 
1079
  def _read_docs(path: Path, text_field: str, skip_bad: bool = False) -> Iterator[dict]:
1080
  """Documents of an input file: .jsonl / .ndjson (one JSON object per line), .json (a JSON array of objects, one
1081
+ object, or JSON Lines), or any other file as one plain-text document; gzip, bzip2 and xz files are decompressed
1082
+ (docs.jsonl.gz is read as .jsonl). Invalid UTF-8 is replaced, with a warning. A bad record (not a JSON object,
1083
+ no string text, an id that cannot be written back) raises BadInput, or with ``skip_bad`` is reported on stderr
1084
+ and skipped. A file that cannot be read at all (see ``_input_problem``) raises BadInput with skippable=False.
1085
+ A null text is kept and scored as an empty document."""
1086
 
1087
  def usable(rec: Any, where: str) -> bool:
1088
  if not isinstance(rec, dict):
 
1091
  problem = f"no {text_field!r} field"
1092
  elif rec[text_field] is not None and not isinstance(rec[text_field], str):
1093
  problem = f"{text_field!r} is a {type(rec[text_field]).__name__}, not a string"
1094
+ elif "id" in rec and (why := _unwritable(rec["id"])):
1095
+ problem = f"the id {why}"
1096
  else:
1097
  return True
1098
  if not skip_bad:
 
1108
  file=sys.stderr)
1109
  return raw.decode("utf-8", errors="replace")
1110
 
1111
+ def read_all() -> bytes:
1112
+ try:
1113
+ with _open_input(path) as f:
1114
+ return f.read()
1115
+ except _READ_ERRORS as e:
1116
+ raise BadInput(f"{path} cannot be read ({e})", skippable=False) from None
1117
+
1118
+ problem = _input_problem(path)
1119
+ if problem:
1120
+ raise BadInput(f"{path} {problem}", skippable=False)
1121
+ suffix = _format_suffix(path)
1122
  if suffix not in (".jsonl", ".ndjson", ".json"):
1123
+ yield {text_field: decode(read_all(), str(path)), "id": path.name}
1124
  return
1125
  if suffix == ".json":
1126
  try:
1127
+ data = json.loads(decode(read_all(), str(path)).lstrip(""))
1128
+ except (ValueError, RecursionError): # JSONDecodeError is a ValueError
1129
+ data = None # not one JSON value (or too large or too deep to read as one): read it as JSON Lines below
1130
  if data is not None:
1131
  for n, rec in enumerate(data if isinstance(data, list) else [data]):
1132
  if usable(rec, f"{path}[{n}]"):
1133
  yield rec
1134
  return
1135
+ n = 0
1136
+ try:
1137
+ with _open_input(path) as f:
1138
+ for n, raw in enumerate(f, 1):
1139
+ where = f"{path}:{n}"
1140
+ line = decode(raw, where)
1141
+ if n == 1:
1142
+ line = line.lstrip("")
1143
+ if not line.strip():
1144
+ continue
1145
+ try:
1146
+ rec = json.loads(line)
1147
+ except (ValueError, RecursionError) as e: # also integers too long to read, and very deep nesting
1148
+ problem = f"invalid JSON ({_json_error(e)})"
1149
+ if not skip_bad:
1150
+ raise BadInput(f"{where}: {problem}") from None
1151
+ print(f"source1: skipped {where}: {problem}", file=sys.stderr)
1152
+ continue
1153
+ if usable(rec, where):
1154
+ yield rec
1155
+ except _READ_ERRORS as e:
1156
+ raise BadInput(f"{path} cannot be read after line {n} ({e})", skippable=False) from None
1157
 
1158
 
1159
  def main(argv: list[str] | None = None) -> int:
 
1161
  p.add_argument("--model", default=str(Path(__file__).resolve().parent),
1162
  help="Source-1 directory or Hugging Face repo id (default: this file's directory)")
1163
  p.add_argument("--input", required=True, help=".jsonl (one document per line), .json (an array of objects) or "
1164
+ "a text file (one document); gzip, bzip2 and xz files are decompressed (docs.jsonl.gz)")
1165
  p.add_argument("--text-field", default="text", help="JSON field holding the text (default: text); "
1166
  "title, url, source_type and code_language fields are used when present")
1167
+ p.add_argument("--output", help="output .jsonl (default: standard output); written to <output>.tmp and renamed "
1168
+ "at the end, so a run that fails leaves an older output as it was")
1169
  p.add_argument("--skip-bad", action="store_true", help="skip (and report on stderr) records that are not valid "
1170
  "JSON objects with a string text, instead of stopping")
1171
  p.add_argument("--revision", help="branch, tag or commit, when --model is a Hugging Face repo id")
 
1175
  p.add_argument("--dtype", choices=DTYPE_CHOICES, default="auto", help="what to compute in; auto (default): "
1176
  "bfloat16 on a GPU with native bfloat16, else float32 (bf16 weights upcast); float16 is not "
1177
  "supported")
1178
+ p.add_argument("--batch-tokens", type=int, help="padded tokens per forward pass (default: "
1179
+ f"{DEFAULT_BATCH_TOKENS_CPU} on a CPU, {DEFAULT_BATCH_TOKENS} on a GPU)")
1180
  p.add_argument("--max-chunks", type=int, default=0, help="score at most N evenly spaced chunks per document")
1181
  p.add_argument("--drop-line", help='"calibrated" (default), "default" (the schema\'s hard filters) or an '
1182
  "expression such as 'toxicity >= 4 or spam_seo >= 3'")
 
1185
  p.add_argument("--no-chunks", action="store_true", help="leave out the per-chunk list of split documents")
1186
  p.add_argument("--group", type=int, default=256, help="documents scored together")
1187
  args = p.parse_args(argv)
1188
+ source = Path(args.input)
1189
+ if not source.is_file():
1190
  p.error(f"--input {args.input}: no such file")
1191
+ problem = _input_problem(source)
1192
+ if problem:
1193
+ p.error(f"--input {args.input} {problem}")
1194
+ if args.batch_tokens is not None and args.batch_tokens < 1:
1195
+ p.error("--batch-tokens must be at least 1")
1196
+ # Checked before the model loads. A regular output file is written as <output>.tmp, which replaces the output
1197
+ # only when the run succeeds; a device or pipe (such as /dev/stdout) is written directly.
1198
+ target = tmp = None
1199
+ if args.output:
1200
+ out_path = Path(args.output)
1201
+ if out_path.is_dir():
1202
+ p.error(f"--output {args.output} is a folder; give a file name")
1203
+ if out_path.exists() and out_path.samefile(source):
1204
+ p.error("--output is the same file as --input; refusing to overwrite it")
1205
+ if not out_path.resolve().parent.is_dir():
1206
+ p.error(f"--output {args.output}: the folder does not exist")
1207
+ if out_path.is_file() or not out_path.exists():
1208
+ target = out_path.resolve()
1209
+ tmp = target.with_name(target.name + ".tmp")
1210
+ if tmp.is_dir() or (tmp.exists() and tmp.samefile(source)):
1211
+ p.error(f"--output {args.output}: the temporary file it is written to first, {tmp}, is "
1212
+ + ("a folder" if tmp.is_dir() else "the --input file"))
1213
 
1214
  t0 = time.time()
1215
+ try:
1216
+ model = Source1.from_pretrained(args.model, device=args.device, dtype=args.dtype, precision=args.precision,
1217
+ drop_line=args.drop_line, apply_offsets=args.apply_offsets,
1218
+ show_url=args.show_url, revision=args.revision)
1219
+ except DropLineError as e:
1220
+ p.error(str(e))
1221
  compute = str(next(model.parameters()).dtype).replace("torch.", "")
1222
  print(f"source1: loaded {model.weights_file} ({model.weights_dtype or '?'} weights) on {model.device}, computing "
1223
  f"in {compute}, in {time.time() - t0:.1f} s; drop line: {model.drop_line.source}", file=sys.stderr)
1224
+ out = open(tmp or args.output, "w", encoding="utf-8") if args.output else sys.stdout
1225
+ done = seen = 0
1226
  t0 = time.time()
1227
 
1228
  def flush(group: list[dict]) -> None:
1229
+ nonlocal done, seen
1230
  for rec, res in zip(group, model.score_batch(group, args.batch_tokens, text_field=args.text_field,
1231
  max_chunks=args.max_chunks)):
1232
+ seen += 1
1233
  if args.no_chunks:
1234
  res.pop("chunks", None)
1235
  if "id" in rec:
1236
  res = {"id": rec["id"], **res}
1237
+ try: # the reader lets no unwritable id through; this keeps a half-written line out of the output
1238
+ line = json.dumps(res, ensure_ascii=False, allow_nan=False)
1239
+ line.encode("utf-8")
1240
+ except (ValueError, RecursionError) as e:
1241
+ where = f"document {seen:,}" + (f" (id {rec['id']!r})" if "id" in rec else "")
1242
+ if not args.skip_bad:
1243
+ raise BadInput(f"{where}: its scores cannot be written as JSON ({e})") from None
1244
+ print(f"source1: skipped {where}: its scores cannot be written as JSON ({e})", file=sys.stderr)
1245
+ continue
1246
+ out.write(line + "\n")
1247
+ done += 1
1248
  print(f"source1: {done:,} documents, {done / max(time.time() - t0, 1e-9):.1f}/s", file=sys.stderr)
1249
 
1250
+ def written() -> str:
1251
+ """What a run that stopped early left behind."""
1252
+ what = f"The {done:,} documents before it were" if done != 1 else "The document before it was"
1253
+ if tmp is None:
1254
+ return f"{what} written"
1255
+ if not done:
1256
+ tmp.unlink(missing_ok=True)
1257
+ return f"Nothing was written, and {args.output} was not changed"
1258
+ return f"{what} written to {tmp}, and {args.output} was not changed"
1259
+
1260
+ docs = _read_docs(source, args.text_field, args.skip_bad)
1261
+ group: list[dict] = []
1262
  try:
 
1263
  try:
1264
+ while True:
1265
+ try:
1266
+ rec = next(docs)
1267
+ except StopIteration:
1268
+ break
1269
+ except Exception:
1270
+ if group:
1271
+ flush(group) # the documents read before a bad record or a read error are still scored
1272
+ raise
1273
  group.append(rec)
1274
  if len(group) >= args.group:
1275
  flush(group)
1276
  group = []
 
1277
  if group:
1278
+ flush(group)
1279
+ finally:
1280
+ if out is not sys.stdout:
1281
+ out.close()
1282
+ except BadInput as e:
1283
+ raise SystemExit(f"source1: {e}. {written()}"
1284
+ + ("; --skip-bad skips bad records" if e.skippable else "")) from None
1285
+ except BaseException as e:
1286
+ if tmp is not None:
1287
+ print(f"source1: stopped by {type(e).__name__}. {written()}", file=sys.stderr)
1288
+ raise
1289
+ if tmp is not None:
1290
+ tmp.replace(target)
1291
  return 0
1292
 
1293