Pre-release review fixes: code robustness, card corrections
Browse files- CREDITS_BOOKS.tsv +27 -25
- EVALUATION.md +59 -29
- NOTICE +21 -14
- README.md +50 -18
- requirements.txt +7 -4
- source1.py +372 -88
CREDITS_BOOKS.tsv
CHANGED
|
@@ -1,8 +1,10 @@
|
|
| 1 |
-
# Source-1: per-work credits for the books in its training and validation splits that are
|
| 2 |
-
# CC0 (2,846 works).
|
| 3 |
-
#
|
| 4 |
-
#
|
| 5 |
-
#
|
|
|
|
|
|
|
| 6 |
# citation_or_note: the citation or attribution that the work's own text asks for, as the work gives it (line
|
| 7 |
# breaks joined, PDF spacing repaired), for FAO books, OpenStax textbooks, Eurydice, JRC and other EU reports, and
|
| 8 |
# some other books and reports (university-press books, Frontiers ebooks, research and project reports); or a
|
|
@@ -107,7 +109,7 @@ DOAB Prácticas lingüísticas heterogéneas: Nuevas perspectivas para el estudi
|
|
| 107 |
DOAB Questioni di donne: Diplomazia informale e reti femminili alla corte dei Savoia-Carignano (XVII secolo) Lurgo, Elisabetta it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://directory.doabooks.org/handle/20.500.12854/134564
|
| 108 |
DOAB Religion og etikk i skole og barnehage Afset, Bente; Redse, Arne no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://directory.doabooks.org/handle/20.500.12854/38379
|
| 109 |
DOAB Réinventer l’art sacré: Le Groupe de Saint-Luc (1919-1945) Noverraz, Camille fr CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://directory.doabooks.org/handle/20.500.12854/152369
|
| 110 |
-
DOAB Sjezd českých právníků 2022
|
| 111 |
DOAB Slovanský literární svět: kontexty a konfrontace III: Motiv domova ve slovanských literaturách Bujnáková, Jana; Cepková Feješová, Zuzana; Derková, Vladimíra; Eniko, Mateja; Heinigová, Lenka; Hrancová, Hana cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80840
|
| 112 |
DOAB Somatopedické simulační techniky a intervence: Metodické texty k projektu MUNI 4.0. Pedagogická fakulta, studijní program Logopedie (Bc.) Opatřilová, Dagmar cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80845
|
| 113 |
DOAB Speciálněpedagogická diagnostika somatopedická: Metodické texty k projektu MUNI 4.0. Pedagogická fakulta, studijní program Logopedie (Bc.) Opatřilová, Dagmar cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80846
|
|
@@ -767,7 +769,7 @@ FAO Combattre la criminalité liée aux forêts en Afrique de l'Ouest Chasi, R.M
|
|
| 767 |
FAO Commission générale des pêches pour la Méditerranée: Rapport de la quarante-huitième session FAO; fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd9408fr Citation requested in the work: FAO. 2026. Commission générale des pêches pour la Méditerranée: Rapport de la quarante-huitième session – Malaga, Espagne, 4-9 novembre 2025. Commission générale des pêches pour la Méditerranée (CGPM) – Rapports de session, n°48. Rome. https://doi.org/10.4060/cd9408fr
|
| 768 |
FAO Compendium of decisions made by the Parties to the FAO Agreement on Port State Measures FAO; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0286en Citation requested in the work: FAO. 2026. Compendium of decisions made by the Parties to the FAO Agreement on Port State Measures. Second edition. Rome. https://doi.org/10.4060/ce0286en
|
| 769 |
FAO Comptes rendus du deuxième Sommet sur la pêche artisanale FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3997fr Citation requested in the work: FAO. 2025. Comptes rendus du deuxième Sommet sur la pêche artisanale, 5-7 juillet 2024, Rome. FAO Comptes rendus des pêches et de l’aquaculture, n° 70. Rome. https://doi.org/10.4060/cd3997fr
|
| 770 |
-
FAO Cостояние мирового рыболовства и аквакультуры – 2026
|
| 771 |
FAO Design of a climate-proof fish buying station Josupeit, H.; Moretti, S.; Sciortino, J.A.; Urbani, R.; van Anrooy, R.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1081en Citation requested in the work: Josupeit, H., Moretti, S., Sciortino, J.A., Urbani, R. & Van Anrooy, R. 2026. Design of a climate-proof fish buying station. FAO Fisheries and Aquaculture Technical Paper, No. 691. Rome, FAO. https://doi.org/10.4060/ce1081en
|
| 772 |
FAO Developing holistic nutrition guidelines and standards for school meals FAO; WFP; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1121en Citation requested in the work: FAO and WFP. 2026. Developing holistic nutrition guidelines and standards for school meals – A global methodology. Rome. https://doi.org/10.4060/ce1121en
|
| 773 |
FAO Developing nutrition-sensitive value chains Andrianarimanana, M.; Galante, A.; Liu, B.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0611en Citation requested in the work: Andrianarimanana, M., Galante, A. & Liu, B. 2026. Developing nutrition‑sensitive value chains – Guidelines for practitioners. Rome, FAO. https://doi.org/10.4060/ce0611en
|
|
@@ -818,7 +820,7 @@ FAO Roles and values of camelids and their products FAO; en CC BY 4.0 https://op
|
|
| 818 |
FAO Scaling up community-base fisheries management in the Pacific Govan, H.; Tuxson, T.; Tauati, M.; Lalavanua, W.; Schwarz, A.; Kinch, J.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1087en Citation requested in the work: Govan, H., Tuxson,T., Tauati, M., Lalavanua, W., Schwarz, A. & Kinch, J. 2026. Scaling up community-based fisheries management in the Pacific – Outlook and prospects for securing sustainable coastal fisheries, livelihoods and ecosystems. Rome, FAO; Noumea, Pacific Community. https://doi.org/10.4060/ce1087en
|
| 819 |
FAO Seed to sip: Arabica coffee production guide for Saudi Arabia Gichimu, B.M.; Ghosh, K.; Alfaifi, B.H.; Alfaifi, K.A.; Al Mutlaq, A.M.; Bustamante Adum, D.; Lubabali, H.A.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0725en Citation requested in the work: Gichimu, B.M, Ghosh, K., Alfaifi, B.H., Alfaifi, K.A., Al Mutlaq, A.M., Bustamante Adum, D. & Lubabali, H.A. 2026. Seed to sip: Arabica coffee production guide for Saudi Arabia. Riyadh, FAO. https://doi.org/10.4060/ce0725en
|
| 820 |
FAO Shaping agrifood systems legislation Rosenbaum, K.L.; Vidar, M.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1406en Citation requested in the work: Rosenbaum, K.L and Vidar, M. 2026. Shaping agrifood systems legislation – Good practices for legal advisors in drafting laws and supporting related processes. Legal Guide No. 5. Rome, FAO. https://doi.org/10.4060/ce1406en
|
| 821 |
-
FAO Soil health and fertilizer Jafari, A.; Le Cotty, T.; Tefft, J.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd9916en Citation requested in the work: Jafari, A., Le Cotty, T. & Tefft, J. 2026. Soil health and fertilizer – Policy and investment prospects in sub-Saharan Africa. Directions in Investment no. 19. Rome, Paris and Brussels, FAO, Agrinatura and the European Union. https://doi.
|
| 822 |
FAO Status of the world's soil resources 2026 FAO; ITPS; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0792en Citation requested in the work: FAO and ITPS. 2026. Status of the world's soil resources 2026. Rome, FAO. https://doi.org/10.4060/ce0792en
|
| 823 |
FAO Strengthening small and medium agroenterprise finance in the Near East and North Africa region Aldredge, H.; Priebe, J.; Zook, D.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0863en Citation requested in the work: Aldredge, H., Priebe, J. & Zook, D. 2026. Strengthening small and medium agroenterprise finance in the Near East and North Africa region: Bridging the finance gap. Rome, FAO. https://doi.org/10.4060/ce0863en
|
| 824 |
FAO Sổ tay hỏi đáp Phùng, T.V.; Tôn, V.Đ.; Sơn, T.H; Phục, N.N.; vi CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0930vi Citation requested in the work: Phùng, T.V., Tôn, V.Đ, Sơn, T.H. and Phục, N.N. 2026. Sổ tay hỏi đáp về th c h nh t t an to n sinh học v xử lý chất thải trong chăn nuôi lợn quy mô vừa v nhỏ. Hà Nội, FAO.
|
|
@@ -832,14 +834,14 @@ FAO Vodič za odgovornu i racionalnu primenu antibiotika kod goveda FAO; sr CC B
|
|
| 832 |
FAO Vodič za odgovornu i racionalnu primenu antibiotika kod ovaca i koza Krnjaić, D.; Savić, B.; Trailović, S.; sr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0604sr Citation requested in the work: Krnjaić, D, Savić, B, Trailović, S. 2026. Vodič za odgovornu i racionalnu primenu antibiotika kod ovaca i koza. Beograd, FAO.
|
| 833 |
FAO Vodič za odgovornu i racionalnu primenu antibiotika kod pasa i mačaka Krnjaić, D.; Resanović, R,; sr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0991sr Citation requested in the work: Krnjaić, D, Resanović, R, Trailović, S. 2026. Vodič za odgovornu i racionalnu primenu antibiotika kod pasa i mačaka . Beograd, FAO.
|
| 834 |
FAO Voices of cooperators: A collective journey through the agrifood systems in the Near East and North Africa FAO; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0708en Citation requested in the work: FAO. 2026. Voices of cooperators: A collective journey through the agrifood systems in the Near East and North Africa region. Cairo. https://doi.org/10.4060/ce0708en
|
| 835 |
-
FAO Wood products in the bioeconomy Reck, B.K.; Johnston, C.; Foong, A.; Gupta, A.; Karpov, A.; Holsten, A.; Misselwitz, P.; Keenan, R.J.; Formenton Cardoso, N.; Walter, S.; Bull, L.; Steel, E.A.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0315en Citation requested in the work: Reck, B.K., Johnston, C., Foong, A., Gupta, A., Karpov, A., Holsten, A., Misselwitz, P., Keenan, R.J., Formenton Cardoso, N., Walter, S., Bull, L., and Steel, E.A. 2026. Wood products in the bioeconomy: Scenario-based assessment of the potential for engineered wood products in climate change mitigation. Rome, FAO. DOI https://
|
| 836 |
FAO Évaluation du projet «Consolidation de la paix dans la zone frontalière du nord-est de la Côte d’Ivoire par le biais d’une gestion participative et planifiée des ressources naturelles» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3911fr Citation requested in the work: FAO. 2025. Évaluation du projet «Consolidation de la paix dans la zone frontalière du nord-est de la Côte d’Ivoire, par le biais d’une gestion participative et planifiée des ressources naturelles» – Code du projet: UNJP/IVC/037/PBF. Série évaluation de projet, 01/2025. Rome. https://doi.org/10.4060/cd3911fr
|
| 837 |
FAO Évaluation du projet «Promouvoir une production de cacao sans déforestation pour réduire les émissions en Côte d’Ivoire» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3686fr Citation requested in the work: FAO. 2025. Évaluation du projet «Promouvoir une production de cacao sans déforestation pour réduire les émissions en Côte d’Ivoire» Rapport de mi-parcous, code du projet: GCP/IVC/609/GCF. Série évaluation de projet, n° 50/2024. Rome. https://doi.org/10.4060/cd3686fr
|
| 838 |
FAO Évaluation du projet «Soutien à l’auto-emploi de la jeunesse rurale, vecteur de paix et de cohésion sociale au Mali» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7607fr Citation requested in the work: FAO. 2025. Évaluation du projet «Soutien à l’auto-emploi de la jeunesse rurale, vecteur de paix et de cohésion sociale au Mali» Code du projet: UNJP/MLI/068/PBF. Série évaluation de projet, n. 26/2025. Rome. https://doi.org/10.4060/cd7607fr
|
| 839 |
FAO Évaluation finale du projet «Revitaliser les agroécosystèmes oasiens à travers une approche durable, intégrée et paysagère dans la région de Drâa Tafilalet» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5845fr Citation requested in the work: FAO. 2025. Évaluation finale du projet «Revitaliser les agroécosystèmes oasiens à travers une approche durable, intégrée et paysagère dans la région de Drâa Tafilalet» - Code du projet: GCP/MOR/046/GFF. Série évaluation de projet, N.° 15/2025. Rome. https://doi.org/10.4060/cd5845fr
|
| 840 |
FAO Évaluation finale du projet «Élimination des polluants organiques persistants et pesticides obsolètes et renforcement de la gestion rationnelle des pesticides au Cameroun» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5148fr Citation requested in the work: FAO. 2025. Évaluation finale du projet «Élimination des polluants organiques persistants et pesticides obsolètes et renforcement de la gestion rationnelle des pesticides au Cameroun» – Code du projet: GCP/CMR/031/GFF, Identifiant FEM: 4641. Série Évaluation de projet, 12/2025 Rome. https://doi.org/10.4060/cd5148fr
|
| 841 |
-
FAO Анализ кооперативного законодательства Российской Федерации и других стран СНГ Клименко, О.И.; Кондракова, И.А.; Мадыгина, О.А.; Горячковская, Ю.М.; Яковлев, В.И.; Касулина, В.В.; ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5542ru Citation requested in the work: Клименко О.И., Кондракова И. А., Мадыгина О. А., Горячковская Ю. М., Яковлев В.И., Касулина В.В. 2025. Анализ кооперативного законодательства Российской Федерации и других стран СНГ. ФАО Законодательное исследование № 119. Рим, ФАО. https://doi.
|
| 842 |
-
FAO Комиссия Кодекс Алиментариус. Руководство по процедуре
|
| 843 |
FAO Положение дел в области продовольствия и сельского хозяйства 2024 ФАО ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd2616ru Citation requested in the work: ФАО. 2024. Положение дел в области продовольствия и сельского хозяйства – 2024. Преобразование агропродовольственных систем с ориентацией на ценностные параметры. Рим. https://doi.org/10.4060/cd2616ru
|
| 844 |
FAO Положение дел в области продовольствия и сельского хозяйства 2025 FAO ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7067ru Citation requested in the work: ФАО. 2025. Положение дел в области продовольствия и сельского хозяйства – 2025. Решение проблемы деградации почв с учетом масштабов землевладений. Рим. https://doi.org/10.4060/cd7067ru
|
| 845 |
FAO Положение дел на рынках сельскохозяйственной продукции – 2024 FAO ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd2144ru Citation requested in the work: ФАО. 2024. Положение дел на рынках сельскохозяйственной продукции – 2024. Торговля и питание: согласованность политики в интересах обеспечения здорового рациона. Рим. https://doi.org/10.4060/cd2144ru
|
|
@@ -848,12 +850,12 @@ FAO Состояние мировых земельных и водных рес
|
|
| 848 |
FAO 全球黑土现状 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc3124zh Citation requested in the work: 。中国北京,中国农业出版社。https://doi.org/10.4060/ 粮农组织。2025。《全球黑土现状》 cc3124zh 20-CP 语学院安排并对翻译的准确性及质量负全部责任。如有出入,应以英文原版为准。
|
| 849 |
FAO 兽药残留对肠道微生物组和人体健康的影响 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5301zh Citation requested in the work: 粮农组织。2025。《兽药残留对肠道微生物组和人体健康的影响——食品安全视角》。中国 北京,中国农业出版社。https://doi.org/10.4060/cc5301zh 20-CP
|
| 850 |
FAO 农药残留对肠道微生物组和人体健康的影响 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5306zh Citation requested in the work: 粮农组织。2025。《农药残留对肠道微生物组和人体健康的影响——食品安全视角》。中国 北京,中国农业出版社。https://doi.org/10.4060/cc5306zh 20-CP
|
| 851 |
-
FAO 卓越数字农业报告
|
| 852 |
FAO 塑料挑战徽章训练手册 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd0922zh Citation requested in the work: 粮农组织。2026。《塑料挑战徽章训练手册》。青年与联合国全球联盟学习和行动系列—— 挑战徽章⑬。中国北京,中国农业出版社。https://doi.org/10.4060/cd0922zh
|
| 853 |
FAO 满足味蕾的养殖水产品 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5140zh Citation requested in the work: 《满足味蕾的养殖水产品——探索十二种地中海与黑海鱼类从海洋至餐桌之 粮农组织。2025。 旅》。中国北京,中国农业出版社。https://doi.org/10.4060/cc5140zh
|
| 854 |
FAO 盐碱土探秘 FAO; IUSS; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc0530zh Citation requested in the work: 粮农组织和国际土壤科学联合会。2025。《盐碱土探秘——全球精选十大儿童科普故事》。 中国北京,中国农业出版社。https://doi.org/10.4060/cc0530zh
|
| 855 |
FAO 社会保护与前瞻行动——保护农业生计 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc7628zh Citation requested in the work: 粮农组织。2026。《社会保护与前瞻行动——保护农业生计》。中国北京,中国农业出版社。 https://doi.org/10.4060/cc7628zh 20-CP 语学院安排并对翻译的准确性及质量负全部责任。如有出入,应以英文原版为准。
|
| 856 |
-
FAO 细胞基食品食用安全解析
|
| 857 |
FAO 能源挑战徽章:生物能源补充训练手册 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd1397zh Citation requested in the work: 粮农组织。2026。《能源挑战徽章:生物能源补充训练手册》。青年与联合国全球联盟学 习和行动系列。中国北京,中国农业出版社。https://doi.org/10.4060/cd1397zh
|
| 858 |
FAO 食品法典委员会程序手册 FAO; WHO; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7978zh Citation requested in the work: 粮农组织和世卫组织。2026。《食品法典委员会程序手册》。第三十一版。罗马。https://doi.org/10.4060/cd4216zh
|
| 859 |
FAO 鱼:知之,烹之,食之 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc1395zh Citation requested in the work: 粮农组织。2025。 《鱼:知之,烹之,食之》。中国北京,中国农业出版社。https://doi.org/10.4060/ cc1395zh
|
|
@@ -933,7 +935,7 @@ Kanripo 周禮疑義擧要 江永 (清) zh CC BY-SA 4.0 https://github.com/kanri
|
|
| 933 |
Kanripo 嘉靖以來首輔傳 王世貞 (明) zh CC BY-SA 4.0 https://github.com/kanripo/KR2g0037
|
| 934 |
Kanripo 埤雅 陸佃 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR1j0011
|
| 935 |
Kanripo 天經惑問 游藝 (清) zh CC BY-SA 4.0 https://github.com/kanripo/KR3f0023
|
| 936 |
-
Kanripo 尚書(正文)
|
| 937 |
Kanripo 尚書全解 林之奇 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR1b0007
|
| 938 |
Kanripo 尚書疑義 馬明衡 (明) zh CC BY-SA 4.0 https://github.com/kanripo/KR1b0039
|
| 939 |
Kanripo 折獄龜鑑 鄭克 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR3c0007
|
|
@@ -1028,13 +1030,13 @@ NDLA NDLA: Yrkesfaglig fordypning (EL-ELE vg1) Albertine Aaberge; Bjørn Dølvin
|
|
| 1028 |
NDLA NDLA: Yrkesliv i barne- og ungdomsarbeiderfag (HS-BUA vg2) Bente Elisabeth Vetland; Camilla Øvstebø ; Cathrine Dunker Furuly; Einar Martin Kålen; Gro Nedberg Grønlid; Guri Bente Hårberg; Hege Nikolaisen; Karl Henrik Aanesen; Kristin Aase; Kristin Sundstrøm; Riborg Anna Ringereide; Rita Enstad-Karlsen, Terranova Media; Siv Stai; Siv Stai ; Tove Engesvik; Vig no CC BY-SA 4.0 https://ndla.no/subject:1:03e810db-3560-47b5-a5f6-e7afe1d0a2d6
|
| 1029 |
NDLA NDLA: Yrkesliv i helsearbeiderfag (HS-HEA vg2) Albertine Aaberge; Aleksandra Krogh; Birgit Flaten; Einar Martin Kålen; Hege Nikolaisen; Hege Nikolaisen ; Helene Grotle; Ingrid Schiefloe Myhre; Johannes Leiknes Nag; Karl Henrik Aanesen; Kathrine Synnøve Karlsen; Lars Sandlie/Høgskolen i Lillehammer; Lene Fossbråten; Marit Smith Sørhøy; NAKU; Odd no CC BY-SA 4.0 https://ndla.no/subject:1:f644f829-4e7a-4e74-a63a-342ef786f68a
|
| 1030 |
OAPEN 100 Cartas para Paulo Freire de quienes pretendemos Enseñar Gárate Vergara, Francisco es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51168
|
| 1031 |
-
OAPEN 20 år med fysikkprestasjoner i fritt fall: Analyser fra TIMSS Advanced og andre internasjonale studier Hole, Arne; Onstad, Torgeir; Hagen, Tor Espen no CC BY (
|
| 1032 |
OAPEN Ai margini del contado: Terra, signoria ed élites locali a Sabbion e nel territorio di Cologna Veneta (secoli XII-XIII) STELLA, Attilio it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60437
|
| 1033 |
OAPEN Bekymringsarbeidet: Politiets forebygging av radikalisering og voldelig ekstremisme Førde, Kristin Engh; Andersen, Arnfinn Jomar; Moum Hellevik, Per no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63675 Citation requested in the work: Førde, K. E, Andersen, A. J. & Hellevik, P. M. (2023). Bekymringsarbeidet. Politiets forebygging av radikalisering og voldelig ekstremisme. Cappelen Damm Akademisk. https://doi.org/10.23865/noasp.185
|
| 1034 |
OAPEN Bewältigung des Scheiterns: Autobiographische Schriften früherer Parteifunktionäre von NSDAP und SED Danner, Hans-Ulrich de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/91039
|
| 1035 |
OAPEN Bologna dopo la pandemia: Impatto territoriale e scenari futuri Castrignanò, Marco; RIMONDI, TOMMASO it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60518
|
| 1036 |
OAPEN Construyendo espacios: la ciudad iberoamericana virreinal: Teoría y estudios de caso Paniagua Pérez, Jesús; Arciello, Daniele es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/51907
|
| 1037 |
-
OAPEN Corsi universitari Fortini, Franco it CC BY (
|
| 1038 |
OAPEN Das kolonisierte Heiligtum: Diskriminierungskritische Perspektiven auf das Verfahren der Musealisierung Balzar, Christoph de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://library.oapen.org/handle/20.500.12657/60744
|
| 1039 |
OAPEN Das sogenannte ‚Königliche Gerichtsbuch‘ – Aufzeichnungen des Michael von Pfullendorf zu den Anfängen des Kammergerichts am römisch-deutschen Königshof (1442 bis 1451): Einführung und Edition Luger, Daniel de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/53427
|
| 1040 |
OAPEN Del Palacio Negro a la Selva Lacandona: Louis Althusser en México Ortega Reyna, Jaime es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63196
|
|
@@ -1045,27 +1047,27 @@ OAPEN Documentación digital y léxico en la traducción e interpretación en lo
|
|
| 1045 |
OAPEN En contra de los impíos. La Actuación de la Buena Prensa Católica en la Arquidiócesis de Santiago, 1906-1936 Loyola, Manuel es CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0/) http://library.oapen.org/handle/20.500.12657/32093
|
| 1046 |
OAPEN Esercizi di ricerca: Dottorato e politiche per la formazione Boffo, Vanna; Togni, Fabio it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62865
|
| 1047 |
OAPEN Estudios Interculturales desde el Sur: procesos, debates y propuestas Samaniego, Mario es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/50252
|
| 1048 |
-
OAPEN Europa: um projecto em construção: Homenagem a David Sassoli Graziani, Michela; Rita, Annabela pt CC BY (
|
| 1049 |
OAPEN Fabulations nocturnes: Écologie, vitalité et opacité dans le cinéma d’Apichatpong Weerasethakul Bordeleau, Érik; Pape, Toni; Rose-Antoinette, Ronald; Szymanski, Adam fr CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) http://library.oapen.org/handle/20.500.12657/31350
|
| 1050 |
OAPEN Fallbuch Asylrecht: Mit Bezügen zum Aufenthaltsrecht Mantel, Johanna; Nachtigall, Rhea; Wasnick, Lars de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0) https://library.oapen.org/handle/20.500.12657/63515
|
| 1051 |
OAPEN Firenze prima degli Uberti: Il ceto dirigente fiorentino nell'XI secolo fra riforme diocesane e affermazione personale e familiare Contessa, Maria Pia it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62867
|
| 1052 |
-
OAPEN Francesco da Barberino al crocevia: Culture, società, bilinguismo Bischetti, Sara; Montefusco, Antonio it CC BY (
|
| 1053 |
OAPEN Fuentes para una Constitución con Poder Indígena Valenzuela, Esteban; Romero, Natacha es CC BY 2.0 (https://creativecommons.org/licenses/by/2.0/) http://library.oapen.org/handle/20.500.12657/32033
|
| 1054 |
OAPEN Förändring och kontinuitet: Reformer inom högre utbildning och forskning 1940–2020 Gribbe, Johan sv CC BY 4.0 (http://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/59842 Citation requested in the work: [Johan Gribbe, 2022, Förändring och kontinuitet: Reformer inom högre utbildning och forskning 1940–2020] Universitetskanslersämbetet. DOI: https://doi.org/10.53340/UKAP-4. Licens: CC-BY 4.0
|
| 1055 |
OAPEN Führt Moral unumgänglich zur Religion?: Zur Kritik der Kantischen Religionsphilosophie bei Jürgen Habermas – eine Entgegnung Langthaler, Rudolf de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51766
|
| 1056 |
OAPEN Geplante Obsoleszenz: Hinter den Kulissen der Produktentwicklung Poppe, Erik; Longmuß, Jörg de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24340
|
| 1057 |
-
OAPEN Gli altri noi: Rom e residenti nella Svizzera italiana: etnografia a mediazione Bizzini, Nadia it CC BY (
|
| 1058 |
OAPEN Gouvernance du secteur de la Sécurité: Leçons des expériences ouest-africaines Bryden, Alan; Chappuis, Fairlie fr CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/32914 Citation requested in the work: Bryden, A et Chappuis, F (dir. publ.) 2015 Gouvernance du secteur de la Sécurité : Leçons des expériences ouest-africaines. London: Ubiquity Press. DOI: http://dx.doi.org/10.5334/bav. Licence: CC-BY 4.0
|
| 1059 |
-
OAPEN Guida al mentoring: Aiutare mentori e allievi ad avere successo Chopra, Vineet; Vaughn, Valerie; Saint, Sanjay it CC BY (
|
| 1060 |
OAPEN Guiden till Spaniensverige: Diaspora, integration och transnationalitet bland svenska föreningar i södra Spanien Olsson, Erik sv CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/27488 Citation requested in the work: Olsson, Erik. 2018. Guiden till Spaniensverige: Diaspora, integration och transnationalitet bland svenska föreningar i södra Spanien. Stockholm: Stockholm University Press. DOI: https://doi.org/10.16993/bao. License: CC-BY
|
| 1061 |
OAPEN Hacia una historia de las tendencias trotskistas después de Trotsky Gaido, Daniel es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58559
|
| 1062 |
OAPEN Handlungsoptionen auf dem Weg in die Gigabit-Gesellschaft: Eine rechtliche Analyse von Konzessions- und Kooperationsmodellen sowie regulatorischer Entflechtungsbestimmungen Toros, Fabian de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/legalcode) https://library.oapen.org/handle/20.500.12657/50285
|
| 1063 |
OAPEN Harpe og sverd: Litteraturhistoriske essay om den norske balladen Solberg, Olav no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/50341
|
| 1064 |
OAPEN Högskolans ansvar: Principer för utveckling av den högre Casson, Andrew sv CC BY 3.0 (https://creativecommons.org/licenses/by/3.0/) http://library.oapen.org/handle/20.500.12657/33044 Citation requested in the work: Casson, A 2015 Högskolans ansvar: Principer för utveckling av den högre utbildningen. London: Ubiquity Press. DOI: http://dx.doi.org/10.5334/bap. License: CC-BY 3.0
|
| 1065 |
-
OAPEN Il Fantasma dell’Io. La massa e l’inconscio mimetico: The Phantom of the Ego: Modernism and the Mimetic Unconscious Lawtoo, Nidesh it CC BY (
|
| 1066 |
OAPEN Il video a 360° nella didattica universitaria: Modelli ed esperienze Ranieri, Maria; Luzzi, Damiana; Cuomo, Stefano it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60444
|
| 1067 |
OAPEN Im Brennpunkt der Wirtschaftspolitik: Innovation, Globalisierung und Klimawandel Keuschnigg, Christian de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/90933
|
| 1068 |
-
OAPEN Immaginare l’altrove nell’epoca dell’Antropocene: Media, confini e cambiamenti climatici CAPPI, VALENTINA it CC BY (
|
| 1069 |
OAPEN Inklusionsorientierte Schulentwicklung: Interdisziplinäre Rückblicke, Einblicke und Ausblicke Frohn, Julia; Bengel, Angelika; Piezunka, Anne; Simon, Toni; Dietze, Torsten de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://library.oapen.org/handle/20.500.12657/60549
|
| 1070 |
OAPEN Innvielse til læreryrket: En analyse av praksislæreres veiledningssamtaler Reier Jensen, Andreas no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24989
|
| 1071 |
OAPEN Interessekonflikter i forskning Ingierd, Helene; Bay-Larsen, Ingrid; Hiis Hauge, Kjellrun no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/25321
|
|
@@ -1088,7 +1090,7 @@ OAPEN La traiettoria storica dell’Etiopia di Meles Zenawi: Fra democrazia rivo
|
|
| 1088 |
OAPEN La trama dell’allegoria: Scritture di ricerca e istanza allegorica nel secondo Novecento italiano Caporiccio, Elisa it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58406
|
| 1089 |
OAPEN La trichera letrada. Intelectuales latinoamericanos y Guerra Fría Alburquerque, Germán es CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0/) http://library.oapen.org/handle/20.500.12657/32089
|
| 1090 |
OAPEN Les normes de prononciation du français: Une étude perceptive panfrancophone Chalier, Marc fr CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51444
|
| 1091 |
-
OAPEN Letras na América Portuguesa: Autores – Textos – Leitores Rodrigues-Moura, Enrique pt CC BY (
|
| 1092 |
OAPEN Lo sguardo territorialista di Leonardo: Il cartografo, l’ingegnere idraulico, il progettista di città e territori Poli, Daniela it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62874
|
| 1093 |
OAPEN L’intervista immaginata: Da genere mediatico a invenzione letteraria GALLERANI, Guido Mattia it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58403
|
| 1094 |
OAPEN L’URSS dentro e fuori: La narrazione italiana del mondo sovietico Traini, Cheti it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60443
|
|
@@ -1102,7 +1104,7 @@ OAPEN Paradigmas y polifuncionalidad: Estudio diacrónico de «preciso»/«preci
|
|
| 1102 |
OAPEN Poéticas espectatoriales en Hispanoamérica y Brasil (1800–1847): Ilustración – emancipación – convivencias excluyentes Fernández, Hans es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/59652
|
| 1103 |
OAPEN Problemáticas étnicas y sociales desde el pensamiento latinoamericano: Temas, Conceptos, Enfoques Kozel, Andrés; Rawicz, Daniela; Devés, Eduardo es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63199
|
| 1104 |
OAPEN PROGETTO STREAMING - STRategiE di mitigazione e gestione dei rischi AMbientalI: casi di studio Nel territorio reGionale Toscano: Azioni locali di sostenibilità: cinque progetti per il futuro del territorio toscano Bartalucci, Chiara; Fagioli, Federico; Giachetti, Andrea; NICCOLAI, ALBERTO; Verdi, Leonardo it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58409
|
| 1105 |
-
OAPEN Pubblicità, educazione e diritto in Kant Perni, Romina it CC BY (
|
| 1106 |
OAPEN Raccontare la Resistenza a scuola: Esperienze e riflessioni Bravi, Luca; Martinelli, Chiara; Oliviero, Stefano it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60445
|
| 1107 |
OAPEN Roher Diamant Dalmatien: Die habsburgische Verwaltung, ihre Probleme und das Land, wie beschrieben von seinem Gouverneur Lilienberg fuer Kaiser Franz I. (1834) Clewing, Konrad de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/legalcode) http://library.oapen.org/handle/20.500.12657/26672 Citation requested in the work: Konrad Clewing (Hg.), Roher Diamant Dalmatien. Die habsburgische Verwaltung, ihre Probleme und das Land, wie beschrieben von seinem Gouverneur Lilienberg für Kaiser Franz I. (1834). München, Berlin, Leipzig, Washington/D.C. 2015
|
| 1108 |
OAPEN Samarbeid om selvhjelp: En antologi om den nye selvhjelpsbevegelsen i Norge Gotaas, Nora; Hatleskog Zeiner, Hilde no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24940
|
|
@@ -1114,7 +1116,7 @@ OAPEN Urbane Transformation durch soziale Innovation: Schlüsselbegriffe und Per
|
|
| 1114 |
OAPEN Wissenschaftskarrieren und Gender Bias: Chancengerechtigkeit an Hochschulen zwischen formellen Vorgaben und informellen Einflüssen Dahmen-Adkins, Jennifer; Wolffram, Andrea de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/90818
|
| 1115 |
OAPEN «Parlare di tutto». Un’idea della critica: Il carteggio Baldacci-Fortini Baldacci, Luigi; Fortini, Franco it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62858
|
| 1116 |
OAPEN Å kjøpe for Norge Langseth, Marius; Similä, Jan Ole no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/49452
|
| 1117 |
-
OAPEN Коммуникативный анализ нехудожественного текста для студентов-магистрантов РКИ Perotto, Monica ru CC BY (
|
| 1118 |
OAPEN Конструкции с опорным глаголом в русском и итальянском языках / Support Verb Constructions. A Russian-Italian Contrastive Analysis MAIKO, TATSIANA ru CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60450
|
| 1119 |
OpenStax Algebra and Trigonometry 2e OpenStax en https://creativecommons.org/licenses/by/4.0/ https://github.com/openstax/osbooks-college-algebra-bundle/blob/4922e46ebc04326979e19392ccc2384cb8b9076c/collections/algebra-and-trigonometry-2e.collection.xml Attribution requested in the work: OpenStax and its content contributors.
|
| 1120 |
OpenStax American Government 4e OpenStax en https://creativecommons.org/licenses/by/4.0/ https://github.com/openstax/osbooks-american-government/blob/5c90dd7907dbf25f42266417e63ca3bb0012f031/collections/american-government-4e.collection.xml Attribution requested in the work: OpenStax and its content contributors.
|
|
|
|
| 1 |
+
# Source-1: per-work credits for the books in the open-books part of its training and validation splits that are
|
| 2 |
+
# not public domain or CC0 (2,846 works). Books and book chapters that came through Common Pile v0.1 (DOAB,
|
| 3 |
+
# Pressbooks, LibreTexts, OER Commons) are credited at collection level in NOTICE 3.5. Each work is credited to
|
| 4 |
+
# the authors and publishers named in it, under the licence recorded for it by the platform it came from (for
|
| 5 |
+
# three works, the IGO licence their own text states; the licence cell says so). Modified: used to train a
|
| 6 |
+
# classifier. No text of these works is distributed with Source-1. This file is part of NOTICE (see NOTICE,
|
| 7 |
+
# part 3).
|
| 8 |
# citation_or_note: the citation or attribution that the work's own text asks for, as the work gives it (line
|
| 9 |
# breaks joined, PDF spacing repaired), for FAO books, OpenStax textbooks, Eurydice, JRC and other EU reports, and
|
| 10 |
# some other books and reports (university-press books, Frontiers ebooks, research and project reports); or a
|
|
|
|
| 109 |
DOAB Questioni di donne: Diplomazia informale e reti femminili alla corte dei Savoia-Carignano (XVII secolo) Lurgo, Elisabetta it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://directory.doabooks.org/handle/20.500.12854/134564
|
| 110 |
DOAB Religion og etikk i skole og barnehage Afset, Bente; Redse, Arne no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://directory.doabooks.org/handle/20.500.12854/38379
|
| 111 |
DOAB Réinventer l’art sacré: Le Groupe de Saint-Luc (1919-1945) Noverraz, Camille fr CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://directory.doabooks.org/handle/20.500.12854/152369
|
| 112 |
+
DOAB Sjezd českých právníků 2022 Jednota českých právníků (collective of authors, no editors recorded) cs CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://directory.doabooks.org/handle/20.500.12854/98695
|
| 113 |
DOAB Slovanský literární svět: kontexty a konfrontace III: Motiv domova ve slovanských literaturách Bujnáková, Jana; Cepková Feješová, Zuzana; Derková, Vladimíra; Eniko, Mateja; Heinigová, Lenka; Hrancová, Hana cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80840
|
| 114 |
DOAB Somatopedické simulační techniky a intervence: Metodické texty k projektu MUNI 4.0. Pedagogická fakulta, studijní program Logopedie (Bc.) Opatřilová, Dagmar cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80845
|
| 115 |
DOAB Speciálněpedagogická diagnostika somatopedická: Metodické texty k projektu MUNI 4.0. Pedagogická fakulta, studijní program Logopedie (Bc.) Opatřilová, Dagmar cs CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://directory.doabooks.org/handle/20.500.12854/80846
|
|
|
|
| 769 |
FAO Commission générale des pêches pour la Méditerranée: Rapport de la quarante-huitième session FAO; fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd9408fr Citation requested in the work: FAO. 2026. Commission générale des pêches pour la Méditerranée: Rapport de la quarante-huitième session – Malaga, Espagne, 4-9 novembre 2025. Commission générale des pêches pour la Méditerranée (CGPM) – Rapports de session, n°48. Rome. https://doi.org/10.4060/cd9408fr
|
| 770 |
FAO Compendium of decisions made by the Parties to the FAO Agreement on Port State Measures FAO; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0286en Citation requested in the work: FAO. 2026. Compendium of decisions made by the Parties to the FAO Agreement on Port State Measures. Second edition. Rome. https://doi.org/10.4060/ce0286en
|
| 771 |
FAO Comptes rendus du deuxième Sommet sur la pêche artisanale FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3997fr Citation requested in the work: FAO. 2025. Comptes rendus du deuxième Sommet sur la pêche artisanale, 5-7 juillet 2024, Rome. FAO Comptes rendus des pêches et de l’aquaculture, n° 70. Rome. https://doi.org/10.4060/cd3997fr
|
| 772 |
+
FAO Cостояние мирового рыболовства и аквакультуры – 2026 ФАО; ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd8357ru Citation requested in the work: "ФАО. 2026. Cостояние мирового рыболовства и аквакультуры – 2026. ""Голубая трансформация"": от замысла к практическим результатам. Рим. https://doi.org/10.4060/cd8357ru"
|
| 773 |
FAO Design of a climate-proof fish buying station Josupeit, H.; Moretti, S.; Sciortino, J.A.; Urbani, R.; van Anrooy, R.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1081en Citation requested in the work: Josupeit, H., Moretti, S., Sciortino, J.A., Urbani, R. & Van Anrooy, R. 2026. Design of a climate-proof fish buying station. FAO Fisheries and Aquaculture Technical Paper, No. 691. Rome, FAO. https://doi.org/10.4060/ce1081en
|
| 774 |
FAO Developing holistic nutrition guidelines and standards for school meals FAO; WFP; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1121en Citation requested in the work: FAO and WFP. 2026. Developing holistic nutrition guidelines and standards for school meals – A global methodology. Rome. https://doi.org/10.4060/ce1121en
|
| 775 |
FAO Developing nutrition-sensitive value chains Andrianarimanana, M.; Galante, A.; Liu, B.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0611en Citation requested in the work: Andrianarimanana, M., Galante, A. & Liu, B. 2026. Developing nutrition‑sensitive value chains – Guidelines for practitioners. Rome, FAO. https://doi.org/10.4060/ce0611en
|
|
|
|
| 820 |
FAO Scaling up community-base fisheries management in the Pacific Govan, H.; Tuxson, T.; Tauati, M.; Lalavanua, W.; Schwarz, A.; Kinch, J.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1087en Citation requested in the work: Govan, H., Tuxson,T., Tauati, M., Lalavanua, W., Schwarz, A. & Kinch, J. 2026. Scaling up community-based fisheries management in the Pacific – Outlook and prospects for securing sustainable coastal fisheries, livelihoods and ecosystems. Rome, FAO; Noumea, Pacific Community. https://doi.org/10.4060/ce1087en
|
| 821 |
FAO Seed to sip: Arabica coffee production guide for Saudi Arabia Gichimu, B.M.; Ghosh, K.; Alfaifi, B.H.; Alfaifi, K.A.; Al Mutlaq, A.M.; Bustamante Adum, D.; Lubabali, H.A.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0725en Citation requested in the work: Gichimu, B.M, Ghosh, K., Alfaifi, B.H., Alfaifi, K.A., Al Mutlaq, A.M., Bustamante Adum, D. & Lubabali, H.A. 2026. Seed to sip: Arabica coffee production guide for Saudi Arabia. Riyadh, FAO. https://doi.org/10.4060/ce0725en
|
| 822 |
FAO Shaping agrifood systems legislation Rosenbaum, K.L.; Vidar, M.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce1406en Citation requested in the work: Rosenbaum, K.L and Vidar, M. 2026. Shaping agrifood systems legislation – Good practices for legal advisors in drafting laws and supporting related processes. Legal Guide No. 5. Rome, FAO. https://doi.org/10.4060/ce1406en
|
| 823 |
+
FAO Soil health and fertilizer Jafari, A.; Le Cotty, T.; Tefft, J.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd9916en Citation requested in the work: Jafari, A., Le Cotty, T. & Tefft, J. 2026. Soil health and fertilizer – Policy and investment prospects in sub-Saharan Africa. Directions in Investment no. 19. Rome, Paris and Brussels, FAO, Agrinatura and the European Union. https://doi.org/10.4060/cd9916en
|
| 824 |
FAO Status of the world's soil resources 2026 FAO; ITPS; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0792en Citation requested in the work: FAO and ITPS. 2026. Status of the world's soil resources 2026. Rome, FAO. https://doi.org/10.4060/ce0792en
|
| 825 |
FAO Strengthening small and medium agroenterprise finance in the Near East and North Africa region Aldredge, H.; Priebe, J.; Zook, D.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0863en Citation requested in the work: Aldredge, H., Priebe, J. & Zook, D. 2026. Strengthening small and medium agroenterprise finance in the Near East and North Africa region: Bridging the finance gap. Rome, FAO. https://doi.org/10.4060/ce0863en
|
| 826 |
FAO Sổ tay hỏi đáp Phùng, T.V.; Tôn, V.Đ.; Sơn, T.H; Phục, N.N.; vi CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0930vi Citation requested in the work: Phùng, T.V., Tôn, V.Đ, Sơn, T.H. and Phục, N.N. 2026. Sổ tay hỏi đáp về th c h nh t t an to n sinh học v xử lý chất thải trong chăn nuôi lợn quy mô vừa v nhỏ. Hà Nội, FAO.
|
|
|
|
| 834 |
FAO Vodič za odgovornu i racionalnu primenu antibiotika kod ovaca i koza Krnjaić, D.; Savić, B.; Trailović, S.; sr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0604sr Citation requested in the work: Krnjaić, D, Savić, B, Trailović, S. 2026. Vodič za odgovornu i racionalnu primenu antibiotika kod ovaca i koza. Beograd, FAO.
|
| 835 |
FAO Vodič za odgovornu i racionalnu primenu antibiotika kod pasa i mačaka Krnjaić, D.; Resanović, R,; sr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0991sr Citation requested in the work: Krnjaić, D, Resanović, R, Trailović, S. 2026. Vodič za odgovornu i racionalnu primenu antibiotika kod pasa i mačaka . Beograd, FAO.
|
| 836 |
FAO Voices of cooperators: A collective journey through the agrifood systems in the Near East and North Africa FAO; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0708en Citation requested in the work: FAO. 2026. Voices of cooperators: A collective journey through the agrifood systems in the Near East and North Africa region. Cairo. https://doi.org/10.4060/ce0708en
|
| 837 |
+
FAO Wood products in the bioeconomy Reck, B.K.; Johnston, C.; Foong, A.; Gupta, A.; Karpov, A.; Holsten, A.; Misselwitz, P.; Keenan, R.J.; Formenton Cardoso, N.; Walter, S.; Bull, L.; Steel, E.A.; en CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/ce0315en Citation requested in the work: Reck, B.K., Johnston, C., Foong, A., Gupta, A., Karpov, A., Holsten, A., Misselwitz, P., Keenan, R.J., Formenton Cardoso, N., Walter, S., Bull, L., and Steel, E.A. 2026. Wood products in the bioeconomy: Scenario-based assessment of the potential for engineered wood products in climate change mitigation. Rome, FAO. DOI https://doi.org/10.4060/ce0315en
|
| 838 |
FAO Évaluation du projet «Consolidation de la paix dans la zone frontalière du nord-est de la Côte d’Ivoire par le biais d’une gestion participative et planifiée des ressources naturelles» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3911fr Citation requested in the work: FAO. 2025. Évaluation du projet «Consolidation de la paix dans la zone frontalière du nord-est de la Côte d’Ivoire, par le biais d’une gestion participative et planifiée des ressources naturelles» – Code du projet: UNJP/IVC/037/PBF. Série évaluation de projet, 01/2025. Rome. https://doi.org/10.4060/cd3911fr
|
| 839 |
FAO Évaluation du projet «Promouvoir une production de cacao sans déforestation pour réduire les émissions en Côte d’Ivoire» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd3686fr Citation requested in the work: FAO. 2025. Évaluation du projet «Promouvoir une production de cacao sans déforestation pour réduire les émissions en Côte d’Ivoire» Rapport de mi-parcous, code du projet: GCP/IVC/609/GCF. Série évaluation de projet, n° 50/2024. Rome. https://doi.org/10.4060/cd3686fr
|
| 840 |
FAO Évaluation du projet «Soutien à l’auto-emploi de la jeunesse rurale, vecteur de paix et de cohésion sociale au Mali» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7607fr Citation requested in the work: FAO. 2025. Évaluation du projet «Soutien à l’auto-emploi de la jeunesse rurale, vecteur de paix et de cohésion sociale au Mali» Code du projet: UNJP/MLI/068/PBF. Série évaluation de projet, n. 26/2025. Rome. https://doi.org/10.4060/cd7607fr
|
| 841 |
FAO Évaluation finale du projet «Revitaliser les agroécosystèmes oasiens à travers une approche durable, intégrée et paysagère dans la région de Drâa Tafilalet» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5845fr Citation requested in the work: FAO. 2025. Évaluation finale du projet «Revitaliser les agroécosystèmes oasiens à travers une approche durable, intégrée et paysagère dans la région de Drâa Tafilalet» - Code du projet: GCP/MOR/046/GFF. Série évaluation de projet, N.° 15/2025. Rome. https://doi.org/10.4060/cd5845fr
|
| 842 |
FAO Évaluation finale du projet «Élimination des polluants organiques persistants et pesticides obsolètes et renforcement de la gestion rationnelle des pesticides au Cameroun» FAO fr CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5148fr Citation requested in the work: FAO. 2025. Évaluation finale du projet «Élimination des polluants organiques persistants et pesticides obsolètes et renforcement de la gestion rationnelle des pesticides au Cameroun» – Code du projet: GCP/CMR/031/GFF, Identifiant FEM: 4641. Série Évaluation de projet, 12/2025 Rome. https://doi.org/10.4060/cd5148fr
|
| 843 |
+
FAO Анализ кооперативного законодательства Российской Федерации и других стран СНГ Клименко, О.И.; Кондракова, И.А.; Мадыгина, О.А.; Горячковская, Ю.М.; Яковлев, В.И.; Касулина, В.В.; ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd5542ru Citation requested in the work: Клименко О.И., Кондракова И. А., Мадыгина О. А., Горячковская Ю. М., Яковлев В.И., Касулина В.В. 2025. Анализ кооперативного законодательства Российской Федерации и других стран СНГ. ФАО Законодательное исследование № 119. Рим, ФАО. https://doi.org/10.4060/cd5542ru
|
| 844 |
+
FAO Комиссия Кодекс Алиментариус. Руководство по процедуре ФАО; ВОЗ; ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7978ru Citation requested in the work: ФАО и ВОЗ. 2026. Комиссия Кодекс Алиментариус. Руководство по процедуре. Тридцать первое издание. Рим. https://doi.org/10.4060/cd7978ru
|
| 845 |
FAO Положение дел в области продовольствия и сельского хозяйства 2024 ФАО ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd2616ru Citation requested in the work: ФАО. 2024. Положение дел в области продовольствия и сельского хозяйства – 2024. Преобразование агропродовольственных систем с ориентацией на ценностные параметры. Рим. https://doi.org/10.4060/cd2616ru
|
| 846 |
FAO Положение дел в области продовольствия и сельского хозяйства 2025 FAO ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7067ru Citation requested in the work: ФАО. 2025. Положение дел в области продовольствия и сельского хозяйства – 2025. Решение проблемы деградации почв с учетом масштабов землевладений. Рим. https://doi.org/10.4060/cd7067ru
|
| 847 |
FAO Положение дел на рынках сельскохозяйственной продукции – 2024 FAO ru CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd2144ru Citation requested in the work: ФАО. 2024. Положение дел на рынках сельскохозяйственной продукции – 2024. Торговля и питание: согласованность политики в интересах обеспечения здорового рациона. Рим. https://doi.org/10.4060/cd2144ru
|
|
|
|
| 850 |
FAO 全球黑土现状 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc3124zh Citation requested in the work: 。中国北京,中国农业出版社。https://doi.org/10.4060/ 粮农组织。2025。《全球黑土现状》 cc3124zh 20-CP 语学院安排并对翻译的准确性及质量负全部责任。如有出入,应以英文原版为准。
|
| 851 |
FAO 兽药残留对肠道微生物组和人体健康的影响 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5301zh Citation requested in the work: 粮农组织。2025。《兽药残留对肠道微生物组和人体健康的影响——食品安全视角》。中国 北京,中国农业出版社。https://doi.org/10.4060/cc5301zh 20-CP
|
| 852 |
FAO 农药残留对肠道微生物组和人体健康的影响 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5306zh Citation requested in the work: 粮农组织。2025。《农药残留对肠道微生物组和人体健康的影响——食品安全视角》。中国 北京,中国农业出版社。https://doi.org/10.4060/cc5306zh 20-CP
|
| 853 |
+
FAO 卓越数字农业报告 FAO; ITU zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc4764zh Citation requested in the work: 粮农组织和国际电信联盟。2025。《卓越数字农业报告——粮农组织与国际电联促进欧洲和中亚数字农业良好做法区域竞赛》。中国北京,中国农业出版社。https://doi.org/10.4060/cc4764zh 20-CPP2021
|
| 854 |
FAO 塑料挑战徽章训练手册 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd0922zh Citation requested in the work: 粮农组织。2026。《塑料挑战徽章训练手册》。青年与联合国全球联盟学习和行动系列—— 挑战徽章⑬。中国北京,中国农业出版社。https://doi.org/10.4060/cd0922zh
|
| 855 |
FAO 满足味蕾的养殖水产品 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc5140zh Citation requested in the work: 《满足味蕾的养殖水产品——探索十二种地中海与黑海鱼类从海洋至餐桌之 粮农组织。2025。 旅》。中国北京,中国农业出版社。https://doi.org/10.4060/cc5140zh
|
| 856 |
FAO 盐碱土探秘 FAO; IUSS; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc0530zh Citation requested in the work: 粮农组织和国际土壤科学联合会。2025。《盐碱土探秘——全球精选十大儿童科普故事》。 中国北京,中国农业出版社。https://doi.org/10.4060/cc0530zh
|
| 857 |
FAO 社会保护与前瞻行动——保护农业生计 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc7628zh Citation requested in the work: 粮农组织。2026。《社会保护与前瞻行动——保护农业生计》。中国北京,中国农业出版社。 https://doi.org/10.4060/cc7628zh 20-CP 语学院安排并对翻译的准确性及质量负全部责任。如有出入,应以英文原版为准。
|
| 858 |
+
FAO 细胞基食品食用安全解析 粮农组织; 世界卫生组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc4855zh Citation requested in the work: 粮农组织和世卫组织。2025。《细胞基食品食用安全解析》 。中国北京,中国农业出版社。 https://doi.org/10.4060/cc4855zh 20-CP。
|
| 859 |
FAO 能源挑战徽章:生物能源补充训练手册 粮农组织; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd1397zh Citation requested in the work: 粮农组织。2026。《能源挑战徽章:生物能源补充训练手册》。青年与联合国全球联盟学 习和行动系列。中国北京,中国农业出版社。https://doi.org/10.4060/cd1397zh
|
| 860 |
FAO 食品法典委员会程序手册 FAO; WHO; zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cd7978zh Citation requested in the work: 粮农组织和世卫组织。2026。《食品法典委员会程序手册》。第三十一版。罗马。https://doi.org/10.4060/cd4216zh
|
| 861 |
FAO 鱼:知之,烹之,食之 FAO zh CC BY 4.0 https://openknowledge.fao.org/handle/20.500.14283/cc1395zh Citation requested in the work: 粮农组织。2025。 《鱼:知之,烹之,食之》。中国北京,中国农业出版社。https://doi.org/10.4060/ cc1395zh
|
|
|
|
| 935 |
Kanripo 嘉靖以來首輔傳 王世貞 (明) zh CC BY-SA 4.0 https://github.com/kanripo/KR2g0037
|
| 936 |
Kanripo 埤雅 陸佃 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR1j0011
|
| 937 |
Kanripo 天經惑問 游藝 (清) zh CC BY-SA 4.0 https://github.com/kanripo/KR3f0023
|
| 938 |
+
Kanripo 尚書(正文) unknown (traditionally attributed to 孔子, Confucius, as compiler) zh CC BY-SA 4.0 https://github.com/kanripo/KR1b0001
|
| 939 |
Kanripo 尚書全解 林之奇 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR1b0007
|
| 940 |
Kanripo 尚書疑義 馬明衡 (明) zh CC BY-SA 4.0 https://github.com/kanripo/KR1b0039
|
| 941 |
Kanripo 折獄龜鑑 鄭克 (宋) zh CC BY-SA 4.0 https://github.com/kanripo/KR3c0007
|
|
|
|
| 1030 |
NDLA NDLA: Yrkesliv i barne- og ungdomsarbeiderfag (HS-BUA vg2) Bente Elisabeth Vetland; Camilla Øvstebø ; Cathrine Dunker Furuly; Einar Martin Kålen; Gro Nedberg Grønlid; Guri Bente Hårberg; Hege Nikolaisen; Karl Henrik Aanesen; Kristin Aase; Kristin Sundstrøm; Riborg Anna Ringereide; Rita Enstad-Karlsen, Terranova Media; Siv Stai; Siv Stai ; Tove Engesvik; Vig no CC BY-SA 4.0 https://ndla.no/subject:1:03e810db-3560-47b5-a5f6-e7afe1d0a2d6
|
| 1031 |
NDLA NDLA: Yrkesliv i helsearbeiderfag (HS-HEA vg2) Albertine Aaberge; Aleksandra Krogh; Birgit Flaten; Einar Martin Kålen; Hege Nikolaisen; Hege Nikolaisen ; Helene Grotle; Ingrid Schiefloe Myhre; Johannes Leiknes Nag; Karl Henrik Aanesen; Kathrine Synnøve Karlsen; Lars Sandlie/Høgskolen i Lillehammer; Lene Fossbråten; Marit Smith Sørhøy; NAKU; Odd no CC BY-SA 4.0 https://ndla.no/subject:1:f644f829-4e7a-4e74-a63a-342ef786f68a
|
| 1032 |
OAPEN 100 Cartas para Paulo Freire de quienes pretendemos Enseñar Gárate Vergara, Francisco es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51168
|
| 1033 |
+
OAPEN 20 år med fysikkprestasjoner i fritt fall: Analyser fra TIMSS Advanced og andre internasjonale studier Hole, Arne; Onstad, Torgeir; Hagen, Tor Espen no CC BY (version not recorded) http://library.oapen.org/handle/20.500.12657/23254
|
| 1034 |
OAPEN Ai margini del contado: Terra, signoria ed élites locali a Sabbion e nel territorio di Cologna Veneta (secoli XII-XIII) STELLA, Attilio it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60437
|
| 1035 |
OAPEN Bekymringsarbeidet: Politiets forebygging av radikalisering og voldelig ekstremisme Førde, Kristin Engh; Andersen, Arnfinn Jomar; Moum Hellevik, Per no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63675 Citation requested in the work: Førde, K. E, Andersen, A. J. & Hellevik, P. M. (2023). Bekymringsarbeidet. Politiets forebygging av radikalisering og voldelig ekstremisme. Cappelen Damm Akademisk. https://doi.org/10.23865/noasp.185
|
| 1036 |
OAPEN Bewältigung des Scheiterns: Autobiographische Schriften früherer Parteifunktionäre von NSDAP und SED Danner, Hans-Ulrich de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/91039
|
| 1037 |
OAPEN Bologna dopo la pandemia: Impatto territoriale e scenari futuri Castrignanò, Marco; RIMONDI, TOMMASO it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60518
|
| 1038 |
OAPEN Construyendo espacios: la ciudad iberoamericana virreinal: Teoría y estudios de caso Paniagua Pérez, Jesús; Arciello, Daniele es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/51907
|
| 1039 |
+
OAPEN Corsi universitari Fortini, Franco it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/89263
|
| 1040 |
OAPEN Das kolonisierte Heiligtum: Diskriminierungskritische Perspektiven auf das Verfahren der Musealisierung Balzar, Christoph de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://library.oapen.org/handle/20.500.12657/60744
|
| 1041 |
OAPEN Das sogenannte ‚Königliche Gerichtsbuch‘ – Aufzeichnungen des Michael von Pfullendorf zu den Anfängen des Kammergerichts am römisch-deutschen Königshof (1442 bis 1451): Einführung und Edition Luger, Daniel de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/53427
|
| 1042 |
OAPEN Del Palacio Negro a la Selva Lacandona: Louis Althusser en México Ortega Reyna, Jaime es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63196
|
|
|
|
| 1047 |
OAPEN En contra de los impíos. La Actuación de la Buena Prensa Católica en la Arquidiócesis de Santiago, 1906-1936 Loyola, Manuel es CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0/) http://library.oapen.org/handle/20.500.12657/32093
|
| 1048 |
OAPEN Esercizi di ricerca: Dottorato e politiche per la formazione Boffo, Vanna; Togni, Fabio it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62865
|
| 1049 |
OAPEN Estudios Interculturales desde el Sur: procesos, debates y propuestas Samaniego, Mario es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/50252
|
| 1050 |
+
OAPEN Europa: um projecto em construção: Homenagem a David Sassoli Graziani, Michela; Rita, Annabela pt CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/62866
|
| 1051 |
OAPEN Fabulations nocturnes: Écologie, vitalité et opacité dans le cinéma d’Apichatpong Weerasethakul Bordeleau, Érik; Pape, Toni; Rose-Antoinette, Ronald; Szymanski, Adam fr CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) http://library.oapen.org/handle/20.500.12657/31350
|
| 1052 |
OAPEN Fallbuch Asylrecht: Mit Bezügen zum Aufenthaltsrecht Mantel, Johanna; Nachtigall, Rhea; Wasnick, Lars de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0) https://library.oapen.org/handle/20.500.12657/63515
|
| 1053 |
OAPEN Firenze prima degli Uberti: Il ceto dirigente fiorentino nell'XI secolo fra riforme diocesane e affermazione personale e familiare Contessa, Maria Pia it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62867
|
| 1054 |
+
OAPEN Francesco da Barberino al crocevia: Culture, società, bilinguismo Bischetti, Sara; Montefusco, Antonio it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/52298
|
| 1055 |
OAPEN Fuentes para una Constitución con Poder Indígena Valenzuela, Esteban; Romero, Natacha es CC BY 2.0 (https://creativecommons.org/licenses/by/2.0/) http://library.oapen.org/handle/20.500.12657/32033
|
| 1056 |
OAPEN Förändring och kontinuitet: Reformer inom högre utbildning och forskning 1940–2020 Gribbe, Johan sv CC BY 4.0 (http://creativecommons.org/licenses/by/4.0) https://library.oapen.org/handle/20.500.12657/59842 Citation requested in the work: [Johan Gribbe, 2022, Förändring och kontinuitet: Reformer inom högre utbildning och forskning 1940–2020] Universitetskanslersämbetet. DOI: https://doi.org/10.53340/UKAP-4. Licens: CC-BY 4.0
|
| 1057 |
OAPEN Führt Moral unumgänglich zur Religion?: Zur Kritik der Kantischen Religionsphilosophie bei Jürgen Habermas – eine Entgegnung Langthaler, Rudolf de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51766
|
| 1058 |
OAPEN Geplante Obsoleszenz: Hinter den Kulissen der Produktentwicklung Poppe, Erik; Longmuß, Jörg de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24340
|
| 1059 |
+
OAPEN Gli altri noi: Rom e residenti nella Svizzera italiana: etnografia a mediazione Bizzini, Nadia it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/41430
|
| 1060 |
OAPEN Gouvernance du secteur de la Sécurité: Leçons des expériences ouest-africaines Bryden, Alan; Chappuis, Fairlie fr CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/32914 Citation requested in the work: Bryden, A et Chappuis, F (dir. publ.) 2015 Gouvernance du secteur de la Sécurité : Leçons des expériences ouest-africaines. London: Ubiquity Press. DOI: http://dx.doi.org/10.5334/bav. Licence: CC-BY 4.0
|
| 1061 |
+
OAPEN Guida al mentoring: Aiutare mentori e allievi ad avere successo Chopra, Vineet; Vaughn, Valerie; Saint, Sanjay it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/96163
|
| 1062 |
OAPEN Guiden till Spaniensverige: Diaspora, integration och transnationalitet bland svenska föreningar i södra Spanien Olsson, Erik sv CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/27488 Citation requested in the work: Olsson, Erik. 2018. Guiden till Spaniensverige: Diaspora, integration och transnationalitet bland svenska föreningar i södra Spanien. Stockholm: Stockholm University Press. DOI: https://doi.org/10.16993/bao. License: CC-BY
|
| 1063 |
OAPEN Hacia una historia de las tendencias trotskistas después de Trotsky Gaido, Daniel es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58559
|
| 1064 |
OAPEN Handlungsoptionen auf dem Weg in die Gigabit-Gesellschaft: Eine rechtliche Analyse von Konzessions- und Kooperationsmodellen sowie regulatorischer Entflechtungsbestimmungen Toros, Fabian de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/legalcode) https://library.oapen.org/handle/20.500.12657/50285
|
| 1065 |
OAPEN Harpe og sverd: Litteraturhistoriske essay om den norske balladen Solberg, Olav no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/50341
|
| 1066 |
OAPEN Högskolans ansvar: Principer för utveckling av den högre Casson, Andrew sv CC BY 3.0 (https://creativecommons.org/licenses/by/3.0/) http://library.oapen.org/handle/20.500.12657/33044 Citation requested in the work: Casson, A 2015 Högskolans ansvar: Principer för utveckling av den högre utbildningen. London: Ubiquity Press. DOI: http://dx.doi.org/10.5334/bap. License: CC-BY 3.0
|
| 1067 |
+
OAPEN Il Fantasma dell’Io. La massa e l’inconscio mimetico: The Phantom of the Ego: Modernism and the Mimetic Unconscious Lawtoo, Nidesh it CC BY (version not recorded) http://library.oapen.org/handle/20.500.12657/25157
|
| 1068 |
OAPEN Il video a 360° nella didattica universitaria: Modelli ed esperienze Ranieri, Maria; Luzzi, Damiana; Cuomo, Stefano it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60444
|
| 1069 |
OAPEN Im Brennpunkt der Wirtschaftspolitik: Innovation, Globalisierung und Klimawandel Keuschnigg, Christian de CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/90933
|
| 1070 |
+
OAPEN Immaginare l’altrove nell’epoca dell’Antropocene: Media, confini e cambiamenti climatici CAPPI, VALENTINA it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/61650
|
| 1071 |
OAPEN Inklusionsorientierte Schulentwicklung: Interdisziplinäre Rückblicke, Einblicke und Ausblicke Frohn, Julia; Bengel, Angelika; Piezunka, Anne; Simon, Toni; Dietze, Torsten de CC BY-SA 4.0 (https://creativecommons.org/licenses/by-sa/4.0/) https://library.oapen.org/handle/20.500.12657/60549
|
| 1072 |
OAPEN Innvielse til læreryrket: En analyse av praksislæreres veiledningssamtaler Reier Jensen, Andreas no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24989
|
| 1073 |
OAPEN Interessekonflikter i forskning Ingierd, Helene; Bay-Larsen, Ingrid; Hiis Hauge, Kjellrun no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/25321
|
|
|
|
| 1090 |
OAPEN La trama dell’allegoria: Scritture di ricerca e istanza allegorica nel secondo Novecento italiano Caporiccio, Elisa it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58406
|
| 1091 |
OAPEN La trichera letrada. Intelectuales latinoamericanos y Guerra Fría Alburquerque, Germán es CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0/) http://library.oapen.org/handle/20.500.12657/32089
|
| 1092 |
OAPEN Les normes de prononciation du français: Une étude perceptive panfrancophone Chalier, Marc fr CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/51444
|
| 1093 |
+
OAPEN Letras na América Portuguesa: Autores – Textos – Leitores Rodrigues-Moura, Enrique pt CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/90323
|
| 1094 |
OAPEN Lo sguardo territorialista di Leonardo: Il cartografo, l’ingegnere idraulico, il progettista di città e territori Poli, Daniela it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62874
|
| 1095 |
OAPEN L’intervista immaginata: Da genere mediatico a invenzione letteraria GALLERANI, Guido Mattia it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58403
|
| 1096 |
OAPEN L’URSS dentro e fuori: La narrazione italiana del mondo sovietico Traini, Cheti it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60443
|
|
|
|
| 1104 |
OAPEN Poéticas espectatoriales en Hispanoamérica y Brasil (1800–1847): Ilustración – emancipación – convivencias excluyentes Fernández, Hans es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/59652
|
| 1105 |
OAPEN Problemáticas étnicas y sociales desde el pensamiento latinoamericano: Temas, Conceptos, Enfoques Kozel, Andrés; Rawicz, Daniela; Devés, Eduardo es CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/63199
|
| 1106 |
OAPEN PROGETTO STREAMING - STRategiE di mitigazione e gestione dei rischi AMbientalI: casi di studio Nel territorio reGionale Toscano: Azioni locali di sostenibilità: cinque progetti per il futuro del territorio toscano Bartalucci, Chiara; Fagioli, Federico; Giachetti, Andrea; NICCOLAI, ALBERTO; Verdi, Leonardo it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/58409
|
| 1107 |
+
OAPEN Pubblicità, educazione e diritto in Kant Perni, Romina it CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/62877
|
| 1108 |
OAPEN Raccontare la Resistenza a scuola: Esperienze e riflessioni Bravi, Luca; Martinelli, Chiara; Oliviero, Stefano it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60445
|
| 1109 |
OAPEN Roher Diamant Dalmatien: Die habsburgische Verwaltung, ihre Probleme und das Land, wie beschrieben von seinem Gouverneur Lilienberg fuer Kaiser Franz I. (1834) Clewing, Konrad de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/legalcode) http://library.oapen.org/handle/20.500.12657/26672 Citation requested in the work: Konrad Clewing (Hg.), Roher Diamant Dalmatien. Die habsburgische Verwaltung, ihre Probleme und das Land, wie beschrieben von seinem Gouverneur Lilienberg für Kaiser Franz I. (1834). München, Berlin, Leipzig, Washington/D.C. 2015
|
| 1110 |
OAPEN Samarbeid om selvhjelp: En antologi om den nye selvhjelpsbevegelsen i Norge Gotaas, Nora; Hatleskog Zeiner, Hilde no CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/) http://library.oapen.org/handle/20.500.12657/24940
|
|
|
|
| 1116 |
OAPEN Wissenschaftskarrieren und Gender Bias: Chancengerechtigkeit an Hochschulen zwischen formellen Vorgaben und informellen Einflüssen Dahmen-Adkins, Jennifer; Wolffram, Andrea de CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/90818
|
| 1117 |
OAPEN «Parlare di tutto». Un’idea della critica: Il carteggio Baldacci-Fortini Baldacci, Luigi; Fortini, Franco it CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/62858
|
| 1118 |
OAPEN Å kjøpe for Norge Langseth, Marius; Similä, Jan Ole no CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/49452
|
| 1119 |
+
OAPEN Коммуникативный анализ нехудожественного текста для студентов-магистрантов РКИ Perotto, Monica ru CC BY (version not recorded) https://library.oapen.org/handle/20.500.12657/89272
|
| 1120 |
OAPEN Конструкции с опорным глаголом в русском и итальянском языках / Support Verb Constructions. A Russian-Italian Contrastive Analysis MAIKO, TATSIANA ru CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) https://library.oapen.org/handle/20.500.12657/60450
|
| 1121 |
OpenStax Algebra and Trigonometry 2e OpenStax en https://creativecommons.org/licenses/by/4.0/ https://github.com/openstax/osbooks-college-algebra-bundle/blob/4922e46ebc04326979e19392ccc2384cb8b9076c/collections/algebra-and-trigonometry-2e.collection.xml Attribution requested in the work: OpenStax and its content contributors.
|
| 1122 |
OpenStax American Government 4e OpenStax en https://creativecommons.org/licenses/by/4.0/ https://github.com/openstax/osbooks-american-government/blob/5c90dd7907dbf25f42266417e63ca3bb0012f031/collections/american-government-4e.collection.xml Attribution requested in the work: OpenStax and its content contributors.
|
EVALUATION.md
CHANGED
|
@@ -1,8 +1,9 @@
|
|
| 1 |
# Source-1 evaluation
|
| 2 |
|
| 3 |
Source-1 was compared with its teacher and 16 public quality scorers on three test sets. An independent proprietary
|
| 4 |
-
LLM grader scored every chunk with Source-1's 13-field rubric
|
| 5 |
-
|
|
|
|
| 6 |
|
| 7 |
## Summary
|
| 8 |
|
|
@@ -38,7 +39,10 @@ were never trained on.
|
|
| 38 |
|
| 39 |
**Home ground.** The held-out documents come from the same kinds of sources as the training data. The public scorers
|
| 40 |
were built for their own definitions of quality, most for educational value, and are not wrong when they disagree with
|
| 41 |
-
this rubric.
|
|
|
|
|
|
|
|
|
|
| 42 |
|
| 43 |
**The three test sets.**
|
| 44 |
|
|
@@ -46,7 +50,7 @@ this rubric.
|
|
| 46 |
|---|---|---|---|---|
|
| 47 |
| held-out set (main result) | 495, one per document | 53 | 64 | documents from Source-1's held-out test split, never trained or calibrated on |
|
| 48 |
| English exam | 414, from 332 documents (413 scored by the teacher) | English | 25 | 200 chunks drawn at random, plus 214 harder cases added on purpose |
|
| 49 |
-
| 12-language exam | 352, one per document | 12 | 46 | web text drawn at random from
|
| 50 |
|
| 51 |
**95% intervals.** Each difference between two models comes with a 95% interval from a paired bootstrap over
|
| 52 |
documents. When the interval excludes zero, chance alone is an unlikely explanation for the difference.
|
|
@@ -67,6 +71,8 @@ grammar its card recommends; the serving engine mainly affects speed, so we clai
|
|
| 67 |
publishes one.
|
| 68 |
- The encoder classifiers ran in Hugging Face transformers as their model cards show, in the precision each card
|
| 69 |
states (bfloat16 where it says so, float32 otherwise).
|
|
|
|
|
|
|
| 70 |
- The two fastText models ran in the fasttext library. They read the whole text with newlines turned into spaces, as
|
| 71 |
the DCLM code does.
|
| 72 |
- EAI-Distill ran in transformers in float32 with greedy decoding (its repository's default settings sample).
|
|
@@ -132,10 +138,12 @@ What grades from the grader's model family did and did not influence:
|
|
| 132 |
- **Held-out set** (the main result): 495 chunks in 53 languages, one per document. All come from Source-1's held-out
|
| 133 |
test split, drawn from the four data stages (web 129, multilingual web 276, conversations/code/synthetic 48, open
|
| 134 |
books 42). The grader drops 64 of them. 159 chunks are English and 29 Chinese; most other languages have 1 to 12
|
| 135 |
-
chunks.
|
| 136 |
-
|
| 137 |
-
|
| 138 |
-
|
|
|
|
|
|
|
| 139 |
- **English exam**: 414 chunks from 332 whole documents split into chunks. 200 chunks from 196 documents were drawn at
|
| 140 |
random from web, wiki, Common Pile, math and code sources (the **random-sample** chunks and documents). 214 chunks
|
| 141 |
from 136 documents were added on purpose to cover harder cases (long documents 57 chunks, academic 50, math 28,
|
|
@@ -145,12 +153,14 @@ What grades from the grader's model family did and did not influence:
|
|
| 145 |
one of the 414 chunks (a random-sample chunk). So the comparisons with the teacher and the public scorers' English
|
| 146 |
rows use the 413 chunks it scored (411 or 412 where a public scorer also lacks one). The 512-token comparison uses
|
| 147 |
all 414.
|
| 148 |
-
- **12-language exam**: 352 chunks, one per document, all
|
| 149 |
-
|
| 150 |
-
|
|
|
|
| 151 |
|
| 152 |
Every exam document was checked against the training data (exact text, URL, title, long-line and shingle matching).
|
| 153 |
-
|
|
|
|
| 154 |
|
| 155 |
#### Near-duplicate check
|
| 156 |
|
|
@@ -162,9 +172,9 @@ each, all short templated texts. 9 in all are covered 20% or more. These chunks
|
|
| 162 |
|
| 163 |
#### Safety filter
|
| 164 |
|
| 165 |
-
A safety filter fixed before
|
| 166 |
-
from Source-1's training, validation and test data. Documents that substantially copy removed text were
|
| 167 |
-
its data as well. Every model is compared on the same filtered chunks.
|
| 168 |
|
| 169 |
</details>
|
| 170 |
|
|
@@ -178,6 +188,10 @@ This file calls either "the grader". That model family also includes one of the
|
|
| 178 |
rubric text. The exams were graded with an older revision of the rubric, from before the rubric settled how to score
|
| 179 |
ads.
|
| 180 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 181 |
The grader returns the 13 fields only. Its reference overall score and keep flag are computed from its scores in the
|
| 182 |
same way as Source-1's: the overall formula of the rubric (see [README.md](README.md#what-you-get)), and keep =
|
| 183 |
false when the rubric's hard filters match (`toxicity >= 4`, `spam_seo >= 4` or `boilerplate >= 4.5`). "The grader's
|
|
@@ -188,10 +202,9 @@ under [How the public scorers were run](#how-the-public-scorers-were-run).
|
|
| 188 |
|
| 189 |
#### Grader self-agreement
|
| 190 |
|
| 191 |
-
The exam grader also graded part of each exam a second time. The two gradings agree at rank 0.
|
| 192 |
-
random-sample
|
| 193 |
-
|
| 194 |
-
closely than two noisy gradings agree with each other.
|
| 195 |
|
| 196 |
</details>
|
| 197 |
|
|
@@ -252,9 +265,10 @@ the grader keeps ([Metric definitions](#metric-definitions)).
|
|
| 252 |
| propella-1 1.7B | 495 | 0.736 | 0.855 | 0.896 | 0.599 (about 38/64) |
|
| 253 |
| JQL-Edu (mean of 3 balanced heads) | 495 | 0.600 | 0.737 | 0.826 | 0.328 (21/64) |
|
| 254 |
|
| 255 |
-
- At this matched rate Source-1 catches 46 of the grader's 64 drops and the teacher 44, a difference within noise
|
| 256 |
-
each model's own drop line (for the teacher, its
|
| 257 |
-
([The drop line](#the-drop-line)). Source-1 and the teacher are also level
|
|
|
|
| 258 |
- The lead over propella-1 4B, the closest public scorer, is at least as large outside English: 0.916 against 0.765
|
| 259 |
on the 335 non-English chunks.
|
| 260 |
|
|
@@ -530,7 +544,7 @@ At the shipped drop line, counted per document (the teacher at its own keep flag
|
|
| 530 |
|---|---|---|---|---|---|
|
| 531 |
| English, all 332 | 24 | 14 (0.583) | 96.7% | 14 (0.583) | 96.1% |
|
| 532 |
| English, the 196 random-sample documents | 13 | 9 (0.692) | 97.4% | 6 (0.462) | 94.9% |
|
| 533 |
-
| 12 languages, all 352 (
|
| 534 |
|
| 535 |
The exams were graded with an older revision of the rubric, before it settled how to score ads (`spam_seo` 3, kept).
|
| 536 |
So part of the gap is rubric drift that affects the teacher and Source-1 alike.
|
|
@@ -707,7 +721,8 @@ the public scorers or the teacher was made under the same conditions, so none is
|
|
| 707 |
Precision and batching: computing in bfloat16, both weight files give the same scores. On these 1,261 chunks, scoring
|
| 708 |
each chunk alone instead of in the default batches moved `overall` by up to 0.04 and a single field by up to 0.10 (3
|
| 709 |
labels changed, no keep decision). float32 differs from bfloat16 by a similar amount (up to 0.03 on `overall` and 0.09
|
| 710 |
-
on a single field; 4 labels and 1 keep decision changed)
|
|
|
|
| 711 |
|
| 712 |
</details>
|
| 713 |
|
|
@@ -716,7 +731,11 @@ on a single field; 4 labels and 1 keep decision changed).
|
|
| 716 |
- **Home ground.** It measures agreement with Source-1's rubric, as applied by graders from one proprietary model
|
| 717 |
family whose grades also steered development. It does not show how Source-1 does on text from other sources, or
|
| 718 |
that filtering with it trains better language models.
|
|
|
|
|
|
|
| 719 |
- **The exams helped choose the teacher**, so they are not independent of the grader.
|
|
|
|
|
|
|
| 720 |
- **It misses about a third of the grader's drops** at the shipped line: it catches 43 of 64 on the held-out set, the
|
| 721 |
teacher, at its own keep flags, 46.
|
| 722 |
- **Like its teacher, it rates some qualities higher than the grader does.** On the English exam, `educational_value`
|
|
@@ -739,7 +758,9 @@ on a single field; 4 labels and 1 keep decision changed).
|
|
| 739 |
<summary>The full text of each limitation</summary>
|
| 740 |
|
| 741 |
- **Home-ground evaluation.** The benchmark measures agreement with Source-1's rubric as applied by independent LLM
|
| 742 |
-
graders from one proprietary model family (one graded the exams, another the held-out set).
|
|
|
|
|
|
|
| 743 |
among other proprietary LLM graders, also steered the rubric revisions, the choice of the teacher and its prompt
|
| 744 |
setup, and the drop-line candidates and floor ([Independence from development](#independence-from-development)).
|
| 745 |
The held-out documents come from the same kinds of sources as the training data. The public scorers were built for
|
|
@@ -747,8 +768,15 @@ on a single field; 4 labels and 1 keep decision changed).
|
|
| 747 |
more generous reading of propella-1 halves its gap to Source-1 ([How propella-1 is read](#how-propella-1-is-read)).
|
| 748 |
The comparison does not show how Source-1 does on text from other sources, or that filtering with Source-1 trains
|
| 749 |
better language models; neither has been tested.
|
|
|
|
|
|
|
|
|
|
| 750 |
- **The exams are not independent of the grader.** They are the samples on which the teacher and its prompt setup were
|
| 751 |
chosen against the exam grader's labels. The held-out set, which played no part in that choice, is the main result.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 752 |
- **It misses about a third of the chunks the grader drops at the shipped line.** On the held-out set the shipped line
|
| 753 |
catches 67.2% of the grader's drops (43/64; interval 55.0% to 77.4%); the teacher catches 71.9% (46/64). 17 of
|
| 754 |
Source-1's 21 misses are also missed by the teacher, so most of what Source-1 misses its teacher misses too. Most
|
|
@@ -842,8 +870,9 @@ cannot be recomputed from this repository alone.
|
|
| 842 |
| code_quality | Not usable code: garbled, minified, obfuscated | Very poor: likely non-functional fragments, no structure, or auto-generated boilerplate | Poor: may work but messy | Acceptable: readable, plausibly correct, minimal docs | Good: clean, idiomatic, documented | Excellent: exemplary, production quality, instructive |
|
| 843 |
| math_quality | Garbled math | Mostly wrong or incoherent | Some correct math, but errors or skipped steps | Generally correct, key steps shown | Correct, clean notation, complete steps | Rigorous and elegant, every step justified |
|
| 844 |
|
| 845 |
-
The rubric anchors and the head layout are in `source1.json`. The teacher's prompt
|
| 846 |
-
are not in `source1.json`
|
|
|
|
| 847 |
|
| 848 |
- Pages whose main purpose is to promote or sell a business, product or service (company "about us" pages, product
|
| 849 |
and landing pages, shop listings, brochures) are ads: format `product_page` and `spam_seo` 3, even when cleanly
|
|
@@ -931,8 +960,9 @@ filtering removed, for evaluation only):
|
|
| 931 |
- Separately, a pre-specified safety filter removed documents from every split, and a rule fixed in advance also
|
| 932 |
removed documents that substantially copy text the safety filter removed (every split).
|
| 933 |
- Book pages that are mostly a table of contents were left out of training (99 chunks).
|
| 934 |
-
- Every exam document was checked against the training data (exact text, URL, title, long-line and shingle matching)
|
| 935 |
-
|
|
|
|
| 936 |
|
| 937 |
#### Labels
|
| 938 |
|
|
|
|
| 1 |
# Source-1 evaluation
|
| 2 |
|
| 3 |
Source-1 was compared with its teacher and 16 public quality scorers on three test sets. An independent proprietary
|
| 4 |
+
LLM grader scored every chunk with Source-1's 13-field rubric: its grades were never trained on, and on the held-out
|
| 5 |
+
set it was given the same instructions as Source-1's teacher. Each number says how closely a model's ranking agrees
|
| 6 |
+
with the grader's, so it measures agreement with this grader applying Source-1's own rubric, on home ground.
|
| 7 |
|
| 8 |
## Summary
|
| 9 |
|
|
|
|
| 39 |
|
| 40 |
**Home ground.** The held-out documents come from the same kinds of sources as the training data. The public scorers
|
| 41 |
were built for their own definitions of quality, most for educational value, and are not wrong when they disagree with
|
| 42 |
+
this rubric. Even a scorer that matched the grader's own `educational_value` scores exactly would reach only 0.874 on
|
| 43 |
+
the held-out set, 0.900 on the English exam and 0.817 on the 12-language exam, so a small part of the gap to the
|
| 44 |
+
educational-value classifiers comes from what this measure asks for;
|
| 45 |
+
[Educational value alone](#educational-value-alone) is the fairer comparison for them.
|
| 46 |
|
| 47 |
**The three test sets.**
|
| 48 |
|
|
|
|
| 50 |
|---|---|---|---|---|
|
| 51 |
| held-out set (main result) | 495, one per document | 53 | 64 | documents from Source-1's held-out test split, never trained or calibrated on |
|
| 52 |
| English exam | 414, from 332 documents (413 scored by the teacher) | English | 25 | 200 chunks drawn at random, plus 214 harder cases added on purpose |
|
| 53 |
+
| 12-language exam | 352, one per document | 12 | 46 | web text drawn from FineWeb-2's test split (blocks of 10 consecutive rows at random positions; Spanish from its first rows), about 30 chunks per language |
|
| 54 |
|
| 55 |
**95% intervals.** Each difference between two models comes with a 95% interval from a paired bootstrap over
|
| 56 |
documents. When the interval excludes zero, chance alone is an unlikely explanation for the difference.
|
|
|
|
| 71 |
publishes one.
|
| 72 |
- The encoder classifiers ran in Hugging Face transformers as their model cards show, in the precision each card
|
| 73 |
states (bfloat16 where it says so, float32 otherwise).
|
| 74 |
+
- Source-1 itself ran in bfloat16, its default. In float32 its rank agreement on the held-out set is unchanged to
|
| 75 |
+
three decimal places.
|
| 76 |
- The two fastText models ran in the fasttext library. They read the whole text with newlines turned into spaces, as
|
| 77 |
the DCLM code does.
|
| 78 |
- EAI-Distill ran in transformers in float32 with greedy decoding (its repository's default settings sample).
|
|
|
|
| 138 |
- **Held-out set** (the main result): 495 chunks in 53 languages, one per document. All come from Source-1's held-out
|
| 139 |
test split, drawn from the four data stages (web 129, multilingual web 276, conversations/code/synthetic 48, open
|
| 140 |
books 42). The grader drops 64 of them. 159 chunks are English and 29 Chinese; most other languages have 1 to 12
|
| 141 |
+
chunks. Chinese was oversampled on purpose: 15 of its 29 chunks were added to the proportional sample, and they hold
|
| 142 |
+
7 of the 64 grader drops. 26 of the 53 languages have 5 or fewer chunks. Source-1 never trained or calibrated on
|
| 143 |
+
these documents. They come from the same kinds of sources as the training data. 42 of the 495 chunks (8.5%) come
|
| 144 |
+
from sources that were later removed from training under the license rules (19 from DCLM-baseline, 16 raw Common
|
| 145 |
+
Crawl pages, and 7 from collections with unreliable or gated license terms). The test split also keeps documents
|
| 146 |
+
that the license filtering removed from training.
|
| 147 |
- **English exam**: 414 chunks from 332 whole documents split into chunks. 200 chunks from 196 documents were drawn at
|
| 148 |
random from web, wiki, Common Pile, math and code sources (the **random-sample** chunks and documents). 214 chunks
|
| 149 |
from 136 documents were added on purpose to cover harder cases (long documents 57 chunks, academic 50, math 28,
|
|
|
|
| 153 |
one of the 414 chunks (a random-sample chunk). So the comparisons with the teacher and the public scorers' English
|
| 154 |
rows use the 413 chunks it scored (411 or 412 where a public scorer also lacks one). The 512-token comparison uses
|
| 155 |
all 414.
|
| 156 |
+
- **12-language exam**: 352 chunks, one per document, all FineWeb-2 web text in 12 languages (ar, bn, de, es, hi,
|
| 157 |
+
ja, ko, ru, sw, th, vi, zh; about 30 each), drawn from FineWeb-2's test split (blocks of 10 consecutive rows at
|
| 158 |
+
random positions; Spanish from its first rows). It has no code or math: 351 of the 352 chunks are plain text by the
|
| 159 |
+
grader's label. The grader drops 46.
|
| 160 |
|
| 161 |
Every exam document was checked against the training data (exact text, URL, title, long-line and shingle matching).
|
| 162 |
+
None has half or more of its text in a training document; the largest share of an exam text found in a training
|
| 163 |
+
document is 41%.
|
| 164 |
|
| 165 |
#### Near-duplicate check
|
| 166 |
|
|
|
|
| 172 |
|
| 173 |
#### Safety filter
|
| 174 |
|
| 175 |
+
A safety filter, with a rule fixed before it was first run, removed a small number of documents from every evaluation
|
| 176 |
+
set and from Source-1's training, validation and test data. Documents that substantially copy removed text were
|
| 177 |
+
removed from its data as well. Every model is compared on the same filtered chunks.
|
| 178 |
|
| 179 |
</details>
|
| 180 |
|
|
|
|
| 188 |
rubric text. The exams were graded with an older revision of the rubric, from before the rubric settled how to score
|
| 189 |
ads.
|
| 190 |
|
| 191 |
+
The held-out grader was given the teacher's own instructions word for word: the 13-field rubric and the special rules
|
| 192 |
+
in [Appendix A](#appendix-a-rubric-anchors), for example that ads are scored `spam_seo` 3 and kept. The public scorers
|
| 193 |
+
follow none of these rules.
|
| 194 |
+
|
| 195 |
The grader returns the 13 fields only. Its reference overall score and keep flag are computed from its scores in the
|
| 196 |
same way as Source-1's: the overall formula of the rubric (see [README.md](README.md#what-you-get)), and keep =
|
| 197 |
false when the rubric's hard filters match (`toxicity >= 4`, `spam_seo >= 4` or `boilerplate >= 4.5`). "The grader's
|
|
|
|
| 202 |
|
| 203 |
#### Grader self-agreement
|
| 204 |
|
| 205 |
+
The exam grader also graded part of each exam a second time. The two gradings agree at rank 0.902 on 97 English
|
| 206 |
+
random-sample chunks and 0.945 on 58 chunks of the 12-language exam. On exactly these chunks Source-1 reaches 0.886
|
| 207 |
+
and 0.892 and the teacher 0.879 and 0.882, below the grader's agreement with itself (within noise for English).
|
|
|
|
| 208 |
|
| 209 |
</details>
|
| 210 |
|
|
|
|
| 265 |
| propella-1 1.7B | 495 | 0.736 | 0.855 | 0.896 | 0.599 (about 38/64) |
|
| 266 |
| JQL-Edu (mean of 3 balanced heads) | 495 | 0.600 | 0.737 | 0.826 | 0.328 (21/64) |
|
| 267 |
|
| 268 |
+
- At this matched rate Source-1 catches 46 of the grader's 64 drops and the teacher 44, a difference within noise
|
| 269 |
+
(without the 15 added Chinese chunks: about 39 and 39 of 57). At each model's own drop line (for the teacher, its
|
| 270 |
+
own keep flags) they catch 43 and 46 ([The drop line](#the-drop-line)). Source-1 and the teacher are also level
|
| 271 |
+
within noise on rank and AUC.
|
| 272 |
- The lead over propella-1 4B, the closest public scorer, is at least as large outside English: 0.916 against 0.765
|
| 273 |
on the 335 non-English chunks.
|
| 274 |
|
|
|
|
| 544 |
|---|---|---|---|---|---|
|
| 545 |
| English, all 332 | 24 | 14 (0.583) | 96.7% | 14 (0.583) | 96.1% |
|
| 546 |
| English, the 196 random-sample documents | 13 | 9 (0.692) | 97.4% | 6 (0.462) | 94.9% |
|
| 547 |
+
| 12 languages, all 352 (none added on purpose) | 46 | 31 (0.674) | 94.3% | 28 (0.609) | 93.5% |
|
| 548 |
|
| 549 |
The exams were graded with an older revision of the rubric, before it settled how to score ads (`spam_seo` 3, kept).
|
| 550 |
So part of the gap is rubric drift that affects the teacher and Source-1 alike.
|
|
|
|
| 721 |
Precision and batching: computing in bfloat16, both weight files give the same scores. On these 1,261 chunks, scoring
|
| 722 |
each chunk alone instead of in the default batches moved `overall` by up to 0.04 and a single field by up to 0.10 (3
|
| 723 |
labels changed, no keep decision). float32 differs from bfloat16 by a similar amount (up to 0.03 on `overall` and 0.09
|
| 724 |
+
on a single field; 4 labels and 1 keep decision changed); computing in float32 with the default bfloat16 weights moved
|
| 725 |
+
`overall` by up to 0.05. In float32, scores do not depend on the batch.
|
| 726 |
|
| 727 |
</details>
|
| 728 |
|
|
|
|
| 731 |
- **Home ground.** It measures agreement with Source-1's rubric, as applied by graders from one proprietary model
|
| 732 |
family whose grades also steered development. It does not show how Source-1 does on text from other sources, or
|
| 733 |
that filtering with it trains better language models.
|
| 734 |
+
- **The numbers are agreement with one grader family, not accuracy.** Another grader applying the same rubric would
|
| 735 |
+
give different values.
|
| 736 |
- **The exams helped choose the teacher**, so they are not independent of the grader.
|
| 737 |
+
- **Agreement is lower among good texts.** Among the chunks the grader keeps, Source-1's rank agreement is 0.87, and in
|
| 738 |
+
the better half of those 0.73.
|
| 739 |
- **It misses about a third of the grader's drops** at the shipped line: it catches 43 of 64 on the held-out set, the
|
| 740 |
teacher, at its own keep flags, 46.
|
| 741 |
- **Like its teacher, it rates some qualities higher than the grader does.** On the English exam, `educational_value`
|
|
|
|
| 758 |
<summary>The full text of each limitation</summary>
|
| 759 |
|
| 760 |
- **Home-ground evaluation.** The benchmark measures agreement with Source-1's rubric as applied by independent LLM
|
| 761 |
+
graders from one proprietary model family (one graded the exams, another the held-out set). They are independent in
|
| 762 |
+
that their grades were never trained on; the held-out grader was given the teacher's own instructions word for word
|
| 763 |
+
([The grader in detail](#the-grader-in-detail)). Grades from that family,
|
| 764 |
among other proprietary LLM graders, also steered the rubric revisions, the choice of the teacher and its prompt
|
| 765 |
setup, and the drop-line candidates and floor ([Independence from development](#independence-from-development)).
|
| 766 |
The held-out documents come from the same kinds of sources as the training data. The public scorers were built for
|
|
|
|
| 768 |
more generous reading of propella-1 halves its gap to Source-1 ([How propella-1 is read](#how-propella-1-is-read)).
|
| 769 |
The comparison does not show how Source-1 does on text from other sources, or that filtering with Source-1 trains
|
| 770 |
better language models; neither has been tested.
|
| 771 |
+
- **The numbers are agreement with one grader family, not accuracy.** Another grader applying the same rubric would
|
| 772 |
+
give different values, and differences of a few hundredths near the top (Source-1 against its teacher) may reflect
|
| 773 |
+
this grader's own habits.
|
| 774 |
- **The exams are not independent of the grader.** They are the samples on which the teacher and its prompt setup were
|
| 775 |
chosen against the exam grader's labels. The held-out set, which played no part in that choice, is the main result.
|
| 776 |
+
- **Agreement is lower among good texts.** About 13% of the held-out chunks are spam, boilerplate or toxic, which are
|
| 777 |
+
easy to tell apart. Among the chunks the grader keeps, Source-1's rank agreement is 0.87, and in the better half of
|
| 778 |
+
those 0.73 (teacher 0.75, propella-1 4B 0.61). Every scorer drops like this on already-filtered text; if you rank
|
| 779 |
+
filtered text, expect the lower figure.
|
| 780 |
- **It misses about a third of the chunks the grader drops at the shipped line.** On the held-out set the shipped line
|
| 781 |
catches 67.2% of the grader's drops (43/64; interval 55.0% to 77.4%); the teacher catches 71.9% (46/64). 17 of
|
| 782 |
Source-1's 21 misses are also missed by the teacher, so most of what Source-1 misses its teacher misses too. Most
|
|
|
|
| 870 |
| code_quality | Not usable code: garbled, minified, obfuscated | Very poor: likely non-functional fragments, no structure, or auto-generated boilerplate | Poor: may work but messy | Acceptable: readable, plausibly correct, minimal docs | Good: clean, idiomatic, documented | Excellent: exemplary, production quality, instructive |
|
| 871 |
| math_quality | Garbled math | Mostly wrong or incoherent | Some correct math, but errors or skipped steps | Generally correct, key steps shown | Correct, clean notation, complete steps | Rigorous and elegant, every step justified |
|
| 872 |
|
| 873 |
+
The rubric anchors and the head layout are in `source1.json`. The teacher's prompt, which the held-out grader also
|
| 874 |
+
received word for word, had a few special rules that are not in `source1.json`. Source-1 was trained on labels that
|
| 875 |
+
follow them, as far as the teacher did:
|
| 876 |
|
| 877 |
- Pages whose main purpose is to promote or sell a business, product or service (company "about us" pages, product
|
| 878 |
and landing pages, shop listings, brochures) are ads: format `product_page` and `spam_seo` 3, even when cleanly
|
|
|
|
| 960 |
- Separately, a pre-specified safety filter removed documents from every split, and a rule fixed in advance also
|
| 961 |
removed documents that substantially copy text the safety filter removed (every split).
|
| 962 |
- Book pages that are mostly a table of contents were left out of training (99 chunks).
|
| 963 |
+
- Every exam document was checked against the training data (exact text, URL, title, long-line and shingle matching).
|
| 964 |
+
None has half or more of its text in a training document; the largest share of an exam text found in a training
|
| 965 |
+
document is 41%.
|
| 966 |
|
| 967 |
#### Labels
|
| 968 |
|
NOTICE
CHANGED
|
@@ -63,10 +63,11 @@ are not part of this release.
|
|
| 63 |
|
| 64 |
Source-1 was evaluated with an independent proprietary LLM grader, used for evaluation
|
| 65 |
only: the grader's outputs were never used as training targets or training data, and
|
| 66 |
-
the shipped drop line was chosen by a fixed rule on the teacher's labels.
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
the
|
|
|
|
| 70 |
development note in README.md and "Independence from development" in EVALUATION.md.
|
| 71 |
|
| 72 |
|
|
@@ -100,13 +101,16 @@ states no open licence; web pages under NC or ND terms, and pages on sites that
|
|
| 100 |
other people's documents; and code whose licence or file header is copyleft,
|
| 101 |
proprietary or not on a permissive allow-list.
|
| 102 |
|
| 103 |
-
Every book
|
| 104 |
-
CREDITS_BOOKS.tsv, which is part of this
|
| 105 |
-
source URL. Where a work's own text asks
|
| 106 |
-
|
| 107 |
-
attribution
|
| 108 |
-
|
| 109 |
-
|
|
|
|
|
|
|
|
|
|
| 110 |
|
| 111 |
License links used below:
|
| 112 |
ODC-By 1.0 https://opendatacommons.org/licenses/by/1-0/
|
|
@@ -372,9 +376,12 @@ licence were removed before training.
|
|
| 372 |
4. CONTACT AND REMOVAL REQUESTS
|
| 373 |
================================================================================
|
| 374 |
|
| 375 |
-
Questions about these credits, corrections, and removal or opt-out requests:
|
| 376 |
-
the Community tab of the model repository,
|
| 377 |
-
https://huggingface.co/msmth/Source-1/discussions.
|
|
|
|
|
|
|
|
|
|
| 378 |
|
| 379 |
|
| 380 |
================================================================================
|
|
|
|
| 63 |
|
| 64 |
Source-1 was evaluated with an independent proprietary LLM grader, used for evaluation
|
| 65 |
only: the grader's outputs were never used as training targets or training data, and
|
| 66 |
+
the shipped drop line was chosen by a fixed rule on the teacher's labels. On the
|
| 67 |
+
held-out set the grader was given the same instructions as the teacher. Grades and
|
| 68 |
+
reviews by proprietary LLMs, among them the grader's model family, did inform some
|
| 69 |
+
design choices (the rubric revision, the choice of teacher and its prompt on the exam
|
| 70 |
+
sets, the candidate drop lines and which collections were filtered out); see the
|
| 71 |
development note in README.md and "Independence from development" in EVALUATION.md.
|
| 72 |
|
| 73 |
|
|
|
|
| 101 |
other people's documents; and code whose licence or file header is copyleft,
|
| 102 |
proprietary or not on a permissive allow-list.
|
| 103 |
|
| 104 |
+
Every book in the open-books part of the training data that is not public domain or
|
| 105 |
+
CC0 is also credited individually in the file CREDITS_BOOKS.tsv, which is part of this
|
| 106 |
+
notice: title, authors, language, licence and source URL. Where a work's own text asks
|
| 107 |
+
to be cited or attributed in a particular way, the file gives that citation or
|
| 108 |
+
attribution: FAO's "Required citation", OpenStax's attribution request, the citations
|
| 109 |
+
requested by Eurydice, JRC and other EU reports, and those of some other books and
|
| 110 |
+
reports (university presses, Frontiers ebooks, research and project reports). The World
|
| 111 |
+
Bank works are credited in Appendix A. Books and book chapters that came through
|
| 112 |
+
Common Pile v0.1 (DOAB, Pressbooks, LibreTexts, OER Commons) are credited at collection
|
| 113 |
+
level in 3.5.
|
| 114 |
|
| 115 |
License links used below:
|
| 116 |
ODC-By 1.0 https://opendatacommons.org/licenses/by/1-0/
|
|
|
|
| 376 |
4. CONTACT AND REMOVAL REQUESTS
|
| 377 |
================================================================================
|
| 378 |
|
| 379 |
+
Questions about these credits, corrections, and removal or opt-out requests: open a
|
| 380 |
+
discussion in the Community tab of the model repository,
|
| 381 |
+
https://huggingface.co/msmth/Source-1/discussions. If your request involves personal
|
| 382 |
+
information, open a discussion without the details and we will arrange a private way to
|
| 383 |
+
reach us. We review every request and, where the content is in our training data,
|
| 384 |
+
exclude it from future versions; published weights cannot be changed.
|
| 385 |
|
| 386 |
|
| 387 |
================================================================================
|
README.md
CHANGED
|
@@ -146,16 +146,19 @@ languages, up to 8,192 tokens at a time. It has 307M parameters and is free to u
|
|
| 146 |
## Highlights
|
| 147 |
|
| 148 |
The numbers are rank agreement with an independent proprietary LLM grader: how closely a model puts texts in the same
|
| 149 |
-
order as the grader does (1.0 means the same order, 0 means no link).
|
| 150 |
-
|
| 151 |
-
|
| 152 |
-
|
| 153 |
-
|
| 154 |
-
|
|
|
|
|
|
|
| 155 |
- **Also ahead when judged on educational value alone**, the thing most public scorers were built for.
|
| 156 |
|
| 157 |
*Caveat: the grader scored with Source-1's own rubric (its scoring guide), so this test plays to Source-1's
|
| 158 |
-
strengths.
|
|
|
|
| 159 |
|
| 160 |
## Quick start
|
| 161 |
|
|
@@ -178,6 +181,16 @@ doc = model.score("Photosynthesis is how plants turn light, water and carbon dio
|
|
| 178 |
print(doc["overall"], doc["keep"]) # a 0-5 score and the keep/drop decision
|
| 179 |
```
|
| 180 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 181 |
<details>
|
| 182 |
<summary>More usage: many texts, precision, options, speed, files</summary>
|
| 183 |
|
|
@@ -206,16 +219,22 @@ python source1.py --model . --input page.txt --device cpu --precision fp32 --dty
|
|
| 206 |
```
|
| 207 |
|
| 208 |
- **Output.** One flat dict per document with the 13 fields, `overall`, `keep` and `drop_reasons`. It also gives the
|
| 209 |
-
length in tokens, the number of chunks (`parts`), whether a chunk was cut, and per-chunk results (`chunks`).
|
|
|
|
|
|
|
| 210 |
- **Precision.** By default it loads `model.safetensors` (bfloat16, 0.6 GB). For the full float32 weights (1.2 GB),
|
| 211 |
remove `--exclude` from the download and pass `precision="fp32"`. It computes in bfloat16 on NVIDIA Ampere or newer
|
| 212 |
GPUs and in float32 elsewhere. float16 is not supported. In bfloat16, scores can shift by up to about 0.04
|
| 213 |
depending on which texts share a batch; pass `dtype="fp32"` or `batch_tokens=1` to avoid that.
|
| 214 |
- **Options.** `drop_line` sets your own drop rule, such as `"toxicity >= 4 or spam_seo >= 3 or boilerplate >= 4.5"`.
|
| 215 |
`max_chunks` scores only some chunks of very long documents, for speed. `revision` pins a Hub version. The command
|
| 216 |
-
line reads JSONL, JSON or plain text, and `python source1.py --help` lists
|
|
|
|
| 217 |
- **Examples and speed.** `examples/` holds six documents and their expected CPU output. A GPU can differ by a few
|
| 218 |
hundredths. Source-1 scores about 30 chunks (50,000 tokens) per second on one RTX 3090 in bfloat16.
|
|
|
|
|
|
|
|
|
|
| 219 |
|
| 220 |
**Files**
|
| 221 |
|
|
@@ -229,12 +248,12 @@ python source1.py --model . --input page.txt --device cpu --precision fp32 --dty
|
|
| 229 |
| `calibration.json` | the drop line and the quality-score offsets, with how they were chosen |
|
| 230 |
| `tokenizer.json`, `tokenizer_config.json` | mmBERT's tokenizer (same vocabulary and merges, re-saved) |
|
| 231 |
| `source1.py` | standalone loader, Python API and command line |
|
| 232 |
-
| `requirements.txt` | `torch`, `transformers`, `safetensors`, `tokenizers` |
|
| 233 |
| `examples/` | six sample inputs (`sample.jsonl`) and their expected command-line output (`expected_output.jsonl`) |
|
| 234 |
| `EVALUATION.md` | the full evaluation, the rubric anchors and the training details |
|
| 235 |
| `images/` | the benchmark chart above |
|
| 236 |
| `LICENSE`, `NOTICE`, `AUTHORS` | license text, third-party notices and credits, authors |
|
| 237 |
-
| `CREDITS_BOOKS.tsv` | per-work credits for the training books that are not public domain (part of `NOTICE`) |
|
| 238 |
|
| 239 |
</details>
|
| 240 |
|
|
@@ -260,9 +279,11 @@ Source-1 returns these 13 fields. Higher is better for quality scores and worse
|
|
| 260 |
scores are `null` when they do not apply. You also get:
|
| 261 |
|
| 262 |
- `overall`: one 0-5 score, the quality scores minus penalties for red flags.
|
| 263 |
-
- `keep`: false (drop) if `toxicity >= 4` or `spam_seo >= 3.5` or `boilerplate >= 4.5`.
|
|
|
|
| 264 |
|
| 265 |
-
Long documents are split into chunks. Each chunk is scored,
|
|
|
|
| 266 |
|
| 267 |
<details>
|
| 268 |
<summary>All fields in detail, the scoring formula, the drop line and long documents</summary>
|
|
@@ -277,8 +298,10 @@ Long documents are split into chunks. Each chunk is scored, and the results are
|
|
| 277 |
**Scores.** Each 0-5 score is a decimal such as 2.73: the model estimates how likely each level from 0 to 5 is, and
|
| 278 |
the score is the average level weighted by those chances. `code_quality` applies when `content_type` is
|
| 279 |
text_with_code or code_only, or `format` is code_file. `math_quality` applies when `content_type` is math_heavy or
|
| 280 |
-
`topic` is math. Otherwise they are `null`.
|
| 281 |
-
|
|
|
|
|
|
|
| 282 |
|
| 283 |
**Formula.**
|
| 284 |
|
|
@@ -299,10 +322,17 @@ fixed rule on the teacher's labels. Stricter lines and what they cost are in
|
|
| 299 |
[EVALUATION.md](EVALUATION.md#the-drop-line). You do not have to use `keep`: ranking by `overall`, or your own rules
|
| 300 |
on single fields, may work better for you.
|
| 301 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 302 |
**Long documents.** A document longer than about 7,800 tokens is split into balanced chunks at natural breaks. Each
|
| 303 |
chunk is scored with a one-line header, as in training (source type, title if given, "Part i of n"). The document gets
|
| 304 |
the label that covers the most tokens and the token-weighted average of each score. `toxicity` takes the maximum, and
|
| 305 |
-
the code and math scores average only the chunks where they apply. `overall` and `keep` are then recomputed
|
|
|
|
|
|
|
| 306 |
|
| 307 |
</details>
|
| 308 |
|
|
@@ -398,8 +428,10 @@ Data sources, filtering, recipe and development details:
|
|
| 398 |
|
| 399 |
## Contact
|
| 400 |
|
| 401 |
-
Questions, corrections and removal requests:
|
| 402 |
-
|
|
|
|
|
|
|
| 403 |
|
| 404 |
## Citation
|
| 405 |
|
|
|
|
| 146 |
## Highlights
|
| 147 |
|
| 148 |
The numbers are rank agreement with an independent proprietary LLM grader: how closely a model puts texts in the same
|
| 149 |
+
order as the grader does (1.0 means the same order, 0 means no link). The grader's grades were never trained on, and on
|
| 150 |
+
the main test set it was given the same instructions as Source-1's teacher.
|
| 151 |
+
|
| 152 |
+
- **Beats each of the 16 public quality scorers we tested, including FineWeb-Edu and propella-1, on every test set it
|
| 153 |
+
was run on** (the 11 English-only scorers were not run on the 12-language exam). On the main test set (495 held-out
|
| 154 |
+
texts in 53 languages): 0.90 vs 0.76 for the best of them, propella-1 4B, a model with 13x more parameters.
|
| 155 |
+
- **Nearly matches its teacher**, the open-weight 27B LLM that labeled its training data, with 1/88 of its parameters:
|
| 156 |
+
0.90 vs 0.91 on the main test set.
|
| 157 |
- **Also ahead when judged on educational value alone**, the thing most public scorers were built for.
|
| 158 |
|
| 159 |
*Caveat: the grader scored with Source-1's own rubric (its scoring guide), so this test plays to Source-1's
|
| 160 |
+
strengths. Even a scorer that matched the grader's educational-value scores exactly would reach only 0.87 here. See
|
| 161 |
+
[Limitations](#limitations) and [EVALUATION.md](EVALUATION.md).*
|
| 162 |
|
| 163 |
## Quick start
|
| 164 |
|
|
|
|
| 181 |
print(doc["overall"], doc["keep"]) # a 0-5 score and the keep/drop decision
|
| 182 |
```
|
| 183 |
|
| 184 |
+
### For AI agents and scripts
|
| 185 |
+
|
| 186 |
+
- **Load it only through `source1.py`.** `pipeline("text-classification", ...)` and `AutoModel...` classes load the
|
| 187 |
+
backbone without Source-1's 13 trained heads and return meaningless `LABEL_0` / `LABEL_1` scores.
|
| 188 |
+
- **Check the setup** with the `examples/sample.jsonl` command above: its output must equal
|
| 189 |
+
`examples/expected_output.jsonl`.
|
| 190 |
+
- **Output:** one JSON object per document (the 13 fields, `overall`, `keep`, `drop_reasons`). **Exit codes:** 0 done;
|
| 191 |
+
2 bad arguments or an unreadable input, before the model loads; 1 a bad record during a run.
|
| 192 |
+
- No `trust_remote_code` and no prompts. Loading from a local folder makes no network calls.
|
| 193 |
+
|
| 194 |
<details>
|
| 195 |
<summary>More usage: many texts, precision, options, speed, files</summary>
|
| 196 |
|
|
|
|
| 219 |
```
|
| 220 |
|
| 221 |
- **Output.** One flat dict per document with the 13 fields, `overall`, `keep` and `drop_reasons`. It also gives the
|
| 222 |
+
length in tokens, the number of chunks (`parts`), whether a chunk was cut, and per-chunk results (`chunks`). Empty
|
| 223 |
+
text (or only spaces, zero-width or control characters) gives `keep` false, `drop_reasons` `["empty text"]` and
|
| 224 |
+
`null` for `overall` and every field, so leave those records out before sorting by `overall`.
|
| 225 |
- **Precision.** By default it loads `model.safetensors` (bfloat16, 0.6 GB). For the full float32 weights (1.2 GB),
|
| 226 |
remove `--exclude` from the download and pass `precision="fp32"`. It computes in bfloat16 on NVIDIA Ampere or newer
|
| 227 |
GPUs and in float32 elsewhere. float16 is not supported. In bfloat16, scores can shift by up to about 0.04
|
| 228 |
depending on which texts share a batch; pass `dtype="fp32"` or `batch_tokens=1` to avoid that.
|
| 229 |
- **Options.** `drop_line` sets your own drop rule, such as `"toxicity >= 4 or spam_seo >= 3 or boilerplate >= 4.5"`.
|
| 230 |
`max_chunks` scores only some chunks of very long documents, for speed. `revision` pins a Hub version. The command
|
| 231 |
+
line reads JSONL, JSON or plain text (also gzip, bzip2 or xz compressed), and `python source1.py --help` lists
|
| 232 |
+
every flag.
|
| 233 |
- **Examples and speed.** `examples/` holds six documents and their expected CPU output. A GPU can differ by a few
|
| 234 |
hundredths. Source-1 scores about 30 chunks (50,000 tokens) per second on one RTX 3090 in bfloat16.
|
| 235 |
+
- **On a CPU.** By default it scores 16,384 tokens per batch there and uses about 3 GB of RAM. On 4 threads it reads
|
| 236 |
+
about 900 tokens per second on typical chunks and about 600 on full-length ones (about 12 seconds per
|
| 237 |
+
7,000-token chunk). Use `--max-chunks` for long documents.
|
| 238 |
|
| 239 |
**Files**
|
| 240 |
|
|
|
|
| 248 |
| `calibration.json` | the drop line and the quality-score offsets, with how they were chosen |
|
| 249 |
| `tokenizer.json`, `tokenizer_config.json` | mmBERT's tokenizer (same vocabulary and merges, re-saved) |
|
| 250 |
| `source1.py` | standalone loader, Python API and command line |
|
| 251 |
+
| `requirements.txt` | `torch`, `transformers`, `safetensors`, `tokenizers`, `huggingface_hub` |
|
| 252 |
| `examples/` | six sample inputs (`sample.jsonl`) and their expected command-line output (`expected_output.jsonl`) |
|
| 253 |
| `EVALUATION.md` | the full evaluation, the rubric anchors and the training details |
|
| 254 |
| `images/` | the benchmark chart above |
|
| 255 |
| `LICENSE`, `NOTICE`, `AUTHORS` | license text, third-party notices and credits, authors |
|
| 256 |
+
| `CREDITS_BOOKS.tsv` | per-work credits for the open-books part of the training data, for books that are not public domain or CC0 (part of `NOTICE`) |
|
| 257 |
|
| 258 |
</details>
|
| 259 |
|
|
|
|
| 279 |
scores are `null` when they do not apply. You also get:
|
| 280 |
|
| 281 |
- `overall`: one 0-5 score, the quality scores minus penalties for red flags.
|
| 282 |
+
- `keep`: false (drop) if `toxicity >= 4` or `spam_seo >= 3.5` or `boilerplate >= 4.5`. It checks only these red
|
| 283 |
+
flags, so also set a threshold on `overall`.
|
| 284 |
|
| 285 |
+
Long documents are split into chunks. Each chunk is scored, the results are combined into one, and `keep` is decided
|
| 286 |
+
on the combined scores (each chunk's own result is in `chunks`).
|
| 287 |
|
| 288 |
<details>
|
| 289 |
<summary>All fields in detail, the scoring formula, the drop line and long documents</summary>
|
|
|
|
| 298 |
**Scores.** Each 0-5 score is a decimal such as 2.73: the model estimates how likely each level from 0 to 5 is, and
|
| 299 |
the score is the average level weighted by those chances. `code_quality` applies when `content_type` is
|
| 300 |
text_with_code or code_only, or `format` is code_file. `math_quality` applies when `content_type` is math_heavy or
|
| 301 |
+
`topic` is math. Otherwise they are `null`. For a split document they average the chunks where they apply, so they can
|
| 302 |
+
be set even when the document's `content_type` is plain_text. What each level means is in `source1.json` and
|
| 303 |
+
[EVALUATION.md Appendix A](EVALUATION.md#appendix-a-rubric-anchors), with the extra rules that the teacher and the
|
| 304 |
+
held-out grader were given.
|
| 305 |
|
| 306 |
**Formula.**
|
| 307 |
|
|
|
|
| 322 |
[EVALUATION.md](EVALUATION.md#the-drop-line). You do not have to use `keep`: ranking by `overall`, or your own rules
|
| 323 |
on single fields, may work better for you.
|
| 324 |
|
| 325 |
+
`keep` applies only this red-flag line, so very short or degenerate text (a single word, an emoji, one letter
|
| 326 |
+
repeated) can still be kept. Combine it with a threshold on `overall`. A drop line can also use `overall`, `tokens`
|
| 327 |
+
and `parts` (the last two for whole documents only), for example
|
| 328 |
+
`"toxicity >= 4 or spam_seo >= 3.5 or boilerplate >= 4.5 or overall < 1"`.
|
| 329 |
+
|
| 330 |
**Long documents.** A document longer than about 7,800 tokens is split into balanced chunks at natural breaks. Each
|
| 331 |
chunk is scored with a one-line header, as in training (source type, title if given, "Part i of n"). The document gets
|
| 332 |
the label that covers the most tokens and the token-weighted average of each score. `toxicity` takes the maximum, and
|
| 333 |
+
the code and math scores average only the chunks where they apply. `overall` and `keep` are then recomputed from these
|
| 334 |
+
combined scores, so a document can be kept even when some of its chunks would be dropped. Each chunk's own scores and
|
| 335 |
+
`keep` are in `chunks`.
|
| 336 |
|
| 337 |
</details>
|
| 338 |
|
|
|
|
| 428 |
|
| 429 |
## Contact
|
| 430 |
|
| 431 |
+
Questions, corrections and removal requests: open a discussion in the
|
| 432 |
+
[Community tab](https://huggingface.co/msmth/Source-1/discussions). If your request involves personal information, open
|
| 433 |
+
a discussion without the details and we will arrange a private way to reach us. We review every request and, where the
|
| 434 |
+
content is in our training data, exclude it from future versions; published weights cannot be changed.
|
| 435 |
|
| 436 |
## Citation
|
| 437 |
|
requirements.txt
CHANGED
|
@@ -1,8 +1,11 @@
|
|
| 1 |
# source1.py needs Python >= 3.10 and these packages. The minimum versions are the oldest this release was tested
|
| 2 |
# with (torch 2.11 and 2.14, transformers 5.17), with both weight files: model.safetensors (bfloat16, the default)
|
| 3 |
-
# and model.fp32.safetensors (float32). Older versions may work but were not tested
|
| 4 |
-
#
|
| 5 |
-
|
| 6 |
-
|
|
|
|
|
|
|
| 7 |
safetensors>=0.8
|
| 8 |
tokenizers>=0.23
|
|
|
|
|
|
| 1 |
# source1.py needs Python >= 3.10 and these packages. The minimum versions are the oldest this release was tested
|
| 2 |
# with (torch 2.11 and 2.14, transformers 5.17), with both weight files: model.safetensors (bfloat16, the default)
|
| 3 |
+
# and model.fp32.safetensors (float32). Older versions may work but were not tested, and the next major versions
|
| 4 |
+
# are left out until they are. huggingface_hub downloads the model when it is loaded by its Hugging Face repo id.
|
| 5 |
+
# No GPU is needed: on a CPU, or a GPU without native bfloat16, the bfloat16 weights are upcast to float32 when
|
| 6 |
+
# loaded.
|
| 7 |
+
torch>=2.11,<3
|
| 8 |
+
transformers>=5.17,<6
|
| 9 |
safetensors>=0.8
|
| 10 |
tokenizers>=0.23
|
| 11 |
+
huggingface_hub>=1.0
|
source1.py
CHANGED
|
@@ -17,6 +17,7 @@ CPU.
|
|
| 17 |
|
| 18 |
python source1.py --input docs.jsonl --output scores.jsonl # one JSON object per line, "text" field
|
| 19 |
python source1.py --input page.txt # a .txt file is one document
|
|
|
|
| 20 |
|
| 21 |
Output: one flat dict per document
|
| 22 |
----------------------------------
|
|
@@ -42,8 +43,8 @@ calibration offsets of calibration.json to the quality scores (about 0.02 at mos
|
|
| 42 |
|
| 43 |
How a document is scored (the same steps the model was trained and evaluated with)
|
| 44 |
------------------------------------------------------------------------------------
|
| 45 |
-
1. ``clean_text``: line endings to "\\n", control characters
|
| 46 |
-
two blank lines in a row.
|
| 47 |
2. ``split_text``: a document longer than 7,808 tokens is split into balanced chunks, each ending at the most
|
| 48 |
natural boundary near its ideal end (headings, then paragraphs, lines, sentences, spaces; definitions in code).
|
| 49 |
3. ``build_input``: each chunk gets a one-line header, a blank line, then the chunk text:
|
|
@@ -51,7 +52,7 @@ How a document is scored (the same steps the model was trained and evaluated wit
|
|
| 51 |
Every training input had a Source line, almost always "dataset record", so that is the default; code files had
|
| 52 |
"<language> source file" (pass ``code_language="Python"``). Title is added when given, Part when the document has
|
| 53 |
more than one chunk. ``url`` is accepted but not shown to the model unless ``show_url=True``: no training input had
|
| 54 |
-
one. On 1,261 held-out benchmark chunks (495 graded held-out chunks from the test split
|
| 55 |
dropping the whole header moved overall by 0.03 on average (at most 0.655), dropping a title by 0.05 on the chunks
|
| 56 |
that had one; adding a URL moved it by up to 0.7 and did not improve its rank agreement with the independent
|
| 57 |
graders of the model card's evaluation.
|
|
@@ -85,13 +86,17 @@ from __future__ import annotations
|
|
| 85 |
|
| 86 |
import argparse
|
| 87 |
import ast
|
|
|
|
|
|
|
| 88 |
import json
|
|
|
|
| 89 |
import math
|
| 90 |
import operator
|
| 91 |
import re
|
| 92 |
import sys
|
| 93 |
import time
|
| 94 |
import unicodedata
|
|
|
|
| 95 |
from bisect import bisect_left
|
| 96 |
from collections.abc import Iterable, Iterator, Mapping
|
| 97 |
from pathlib import Path
|
|
@@ -105,7 +110,8 @@ __version__ = "1.0.0"
|
|
| 105 |
MAX_LENGTH = 8192 # tokens per model input, <bos> and <eos> included
|
| 106 |
CHUNK_TOKENS = 7808 # document tokens per chunk: 8,192 minus 384 kept for the header (as in training)
|
| 107 |
DEFAULT_SOURCE = "dataset record"
|
| 108 |
-
DEFAULT_BATCH_TOKENS = 65536 # padded tokens per forward pass
|
|
|
|
| 109 |
LEVELS = (0, 1, 2, 3, 4, 5)
|
| 110 |
# The backbone weights of each precision: bfloat16 (the default) and the full-precision float32 copy. The fp32 file
|
| 111 |
# follows transformers' variant naming (model.<variant>.safetensors), so AutoModel loads it with variant="fp32".
|
|
@@ -126,7 +132,9 @@ def hub_files(precision: str = DEFAULT_PRECISION) -> tuple[str, ...]:
|
|
| 126 |
|
| 127 |
|
| 128 |
_JUNK = re.compile("[\x00-\x08\x0b\x0e-\x1f\x7f\ud800-\udfff\ufeff\u200b\ufffe\uffff]")
|
| 129 |
-
|
|
|
|
|
|
|
| 130 |
_BLANK_LINES = re.compile(r"\n{4,}")
|
| 131 |
_SPACE_RUN = re.compile(r"[ \t]{2,}")
|
| 132 |
_BLANK_RUN = re.compile(r"\n(?:[ \t]*\n){3,}")
|
|
@@ -134,8 +142,9 @@ _BLANK_RUN = re.compile(r"\n(?:[ \t]*\n){3,}")
|
|
| 134 |
|
| 135 |
def clean_text(text: str) -> str:
|
| 136 |
"""The document as the scorer sees it before chunking: "\\n" line endings (a form feed counts as a paragraph
|
| 137 |
-
break), control characters, zero-width spaces and byte-order
|
| 138 |
-
|
|
|
|
| 139 |
text = text.replace("\r\n", "\n").replace("\r", "\n").replace("\x0c", "\n\n")
|
| 140 |
text = _JUNK.sub("", text)
|
| 141 |
text = unicodedata.normalize("NFC", text)
|
|
@@ -296,25 +305,107 @@ _CMP = {ast.Eq: operator.eq, ast.NotEq: operator.ne, ast.Lt: operator.lt, ast.Lt
|
|
| 296 |
ast.Gt: operator.gt, ast.GtE: operator.ge}
|
| 297 |
|
| 298 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 299 |
class DropLine:
|
| 300 |
-
"""A drop line such as ``toxicity >= 4 or spam_seo >= 3.5 or boilerplate >= 4.5``:
|
| 301 |
-
|
| 302 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 303 |
|
| 304 |
_NODES = (ast.Expression, ast.BoolOp, ast.And, ast.Or, ast.UnaryOp, ast.Not, ast.USub, ast.Compare, ast.Name,
|
| 305 |
ast.Load, ast.Constant, *_CMP)
|
| 306 |
|
| 307 |
-
def __init__(self, source: str, names: Iterable[str]):
|
| 308 |
self.source = source.strip()
|
| 309 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 310 |
for node in ast.walk(tree):
|
| 311 |
if not isinstance(node, self._NODES):
|
| 312 |
-
raise
|
| 313 |
-
if isinstance(node, ast.Name) and node.id not in
|
| 314 |
-
raise
|
| 315 |
self.tree = tree.body
|
|
|
|
|
|
|
|
|
|
|
|
|
| 316 |
top_or = isinstance(self.tree, ast.BoolOp) and isinstance(self.tree.op, ast.Or)
|
| 317 |
self.terms = list(self.tree.values) if top_or else [self.tree]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 318 |
|
| 319 |
def reasons(self, env: dict) -> list[str]:
|
| 320 |
"""The terms of the line that match ``env`` (each ``or`` branch on its own); [] means keep."""
|
|
@@ -341,16 +432,11 @@ class DropLine:
|
|
| 341 |
left = self._eval(node.left, env)
|
| 342 |
for op, comp in zip(node.ops, node.comparators):
|
| 343 |
right = self._eval(comp, env)
|
| 344 |
-
if
|
| 345 |
-
return False
|
| 346 |
-
try:
|
| 347 |
-
if not _CMP[type(op)](left, right):
|
| 348 |
-
return False
|
| 349 |
-
except TypeError:
|
| 350 |
-
return False
|
| 351 |
left = right
|
| 352 |
return True
|
| 353 |
-
raise
|
| 354 |
|
| 355 |
|
| 356 |
def _r3(x: float | None) -> float | None:
|
|
@@ -567,6 +653,23 @@ def _load_calibration(path: Path, drop_line: str | None, apply_offsets: bool) ->
|
|
| 567 |
return calibration
|
| 568 |
|
| 569 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 570 |
class Source1(nn.Module):
|
| 571 |
"""Source-1: the mmBERT-base encoder, mean pooling and one linear head per field. Build it with
|
| 572 |
``Source1.from_pretrained``; score documents with ``score`` / ``score_batch``."""
|
|
@@ -590,15 +693,7 @@ class Source1(nn.Module):
|
|
| 590 |
self.labels = {name: list(spec["values"]) for name, spec in self.schema["labels"].items()}
|
| 591 |
self.fields = list(self.labels) + [n for g in ("quality", "red_flags", "gated") for n in self.schema.get(g, {})]
|
| 592 |
self.calibration = calibration or {}
|
| 593 |
-
|
| 594 |
-
if drop_line in (None, "calibrated"):
|
| 595 |
-
drop_line = (self.calibration.get("drop_line") or {}).get("line")
|
| 596 |
-
if not drop_line:
|
| 597 |
-
raise ValueError("no calibrated drop line (calibration.json missing or incomplete); pass "
|
| 598 |
-
"drop_line='default' or your own drop line")
|
| 599 |
-
elif drop_line == "default":
|
| 600 |
-
drop_line = default_line
|
| 601 |
-
self.drop_line = DropLine(drop_line, set(self.fields) | {"overall", "tokens", "parts"})
|
| 602 |
self.offsets = {k: float(v["offset"]) for k, v in (self.calibration.get("offsets") or {}).items()}
|
| 603 |
self.apply_offsets = apply_offsets
|
| 604 |
self.show_url = show_url
|
|
@@ -658,11 +753,21 @@ class Source1(nn.Module):
|
|
| 658 |
raise FileNotFoundError(f"{path} lacks {', '.join(missing)}")
|
| 659 |
config = json.loads((path / "source1.json").read_text(encoding="utf-8"))
|
| 660 |
calibration = _load_calibration(path, drop_line, apply_offsets)
|
|
|
|
| 661 |
tok = Tokenizer.from_file(str(path / "tokenizer.json"))
|
| 662 |
tok.no_truncation()
|
| 663 |
tok.no_padding()
|
| 664 |
variant = {"variant": "fp32"} if precision == "fp32" else {}
|
| 665 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 666 |
model = cls(backbone, tok, config, calibration, drop_line, apply_offsets, show_url)
|
| 667 |
model.heads.load_state_dict(load_file(str(path / "heads.safetensors")))
|
| 668 |
model.heads.to(dtype)
|
|
@@ -754,12 +859,19 @@ class Source1(nn.Module):
|
|
| 754 |
rec["drop_reasons"] = reasons
|
| 755 |
return rec
|
| 756 |
|
| 757 |
-
def
|
|
|
|
|
|
|
|
|
|
|
|
|
| 758 |
"""Score ready-made model inputs (``build_input`` output: header, blank line, chunk text), one dict per
|
| 759 |
input with the 13 fields, overall, keep, drop_reasons, input_tokens and truncated.
|
| 760 |
|
| 761 |
-
Inputs are sorted by length and batched with at most ``batch_tokens`` padded tokens per forward pass
|
| 762 |
-
a batch that runs out of GPU memory is split in half and
|
|
|
|
|
|
|
|
|
|
| 763 |
encoded = self.encode(inputs)
|
| 764 |
ids = [e[0] for e in encoded]
|
| 765 |
results: list[dict | None] = [None] * len(ids)
|
|
@@ -789,13 +901,13 @@ class Source1(nn.Module):
|
|
| 789 |
run(batch)
|
| 790 |
return results # type: ignore[return-value]
|
| 791 |
|
| 792 |
-
def score_batch(self, docs: Iterable[str | bytes | dict], batch_tokens: int =
|
| 793 |
text_field: str = "text", max_chunks: int = 0) -> list[dict]:
|
| 794 |
"""Score a list of documents: strings (bytes are decoded as UTF-8), or dicts with the text under
|
| 795 |
``text_field`` and optionally ``title``, ``url``, ``source_type`` and ``code_language`` (see ``score``).
|
| 796 |
A None text is scored as an empty document (keep False, drop_reasons ["empty text"]). Chunks of all
|
| 797 |
documents are batched together. ``max_chunks`` > 0 scores only that many evenly spaced chunks of a long
|
| 798 |
-
document (the Part numbers still count every chunk).
|
| 799 |
|
| 800 |
In bfloat16 a document's scores can shift slightly (up to about 0.04 on overall, 0.10 on a single field)
|
| 801 |
depending on which other documents share its batch, because the batch shape changes the kernels' rounding;
|
|
@@ -831,7 +943,7 @@ class Source1(nn.Module):
|
|
| 831 |
|
| 832 |
def score(self, text: str | bytes, title: str | None = None, url: str | None = None, *,
|
| 833 |
source_type: str | None = None, code_language: str | None = None, max_chunks: int = 0,
|
| 834 |
-
batch_tokens: int =
|
| 835 |
"""Score one document (a str; bytes are decoded as UTF-8; anything else raises TypeError).
|
| 836 |
|
| 837 |
title: shown to the model in the header when given (as in training, where about a quarter of inputs had one).
|
|
@@ -869,14 +981,108 @@ class Source1(nn.Module):
|
|
| 869 |
|
| 870 |
|
| 871 |
class BadInput(ValueError):
|
| 872 |
-
"""A record of the input file that cannot be scored (the message starts with file:line)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 873 |
|
| 874 |
|
| 875 |
def _read_docs(path: Path, text_field: str, skip_bad: bool = False) -> Iterator[dict]:
|
| 876 |
"""Documents of an input file: .jsonl / .ndjson (one JSON object per line), .json (a JSON array of objects, one
|
| 877 |
-
object, or JSON Lines), or any other file as one plain-text document
|
| 878 |
-
|
| 879 |
-
|
|
|
|
|
|
|
| 880 |
|
| 881 |
def usable(rec: Any, where: str) -> bool:
|
| 882 |
if not isinstance(rec, dict):
|
|
@@ -885,6 +1091,8 @@ def _read_docs(path: Path, text_field: str, skip_bad: bool = False) -> Iterator[
|
|
| 885 |
problem = f"no {text_field!r} field"
|
| 886 |
elif rec[text_field] is not None and not isinstance(rec[text_field], str):
|
| 887 |
problem = f"{text_field!r} is a {type(rec[text_field]).__name__}, not a string"
|
|
|
|
|
|
|
| 888 |
else:
|
| 889 |
return True
|
| 890 |
if not skip_bad:
|
|
@@ -900,37 +1108,52 @@ def _read_docs(path: Path, text_field: str, skip_bad: bool = False) -> Iterator[
|
|
| 900 |
file=sys.stderr)
|
| 901 |
return raw.decode("utf-8", errors="replace")
|
| 902 |
|
| 903 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 904 |
if suffix not in (".jsonl", ".ndjson", ".json"):
|
| 905 |
-
yield {text_field:
|
| 906 |
return
|
| 907 |
if suffix == ".json":
|
| 908 |
try:
|
| 909 |
-
data = json.loads(decode(
|
| 910 |
-
except
|
| 911 |
-
data = None # not
|
| 912 |
if data is not None:
|
| 913 |
for n, rec in enumerate(data if isinstance(data, list) else [data]):
|
| 914 |
if usable(rec, f"{path}[{n}]"):
|
| 915 |
yield rec
|
| 916 |
return
|
| 917 |
-
|
| 918 |
-
|
| 919 |
-
|
| 920 |
-
|
| 921 |
-
|
| 922 |
-
line =
|
| 923 |
-
|
| 924 |
-
|
| 925 |
-
|
| 926 |
-
|
| 927 |
-
|
| 928 |
-
|
| 929 |
-
|
| 930 |
-
|
| 931 |
-
|
| 932 |
-
|
| 933 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 934 |
|
| 935 |
|
| 936 |
def main(argv: list[str] | None = None) -> int:
|
|
@@ -938,10 +1161,11 @@ def main(argv: list[str] | None = None) -> int:
|
|
| 938 |
p.add_argument("--model", default=str(Path(__file__).resolve().parent),
|
| 939 |
help="Source-1 directory or Hugging Face repo id (default: this file's directory)")
|
| 940 |
p.add_argument("--input", required=True, help=".jsonl (one document per line), .json (an array of objects) or "
|
| 941 |
-
"a text file (one document)")
|
| 942 |
p.add_argument("--text-field", default="text", help="JSON field holding the text (default: text); "
|
| 943 |
"title, url, source_type and code_language fields are used when present")
|
| 944 |
-
p.add_argument("--output", help="output .jsonl (default: standard output)"
|
|
|
|
| 945 |
p.add_argument("--skip-bad", action="store_true", help="skip (and report on stderr) records that are not valid "
|
| 946 |
"JSON objects with a string text, instead of stopping")
|
| 947 |
p.add_argument("--revision", help="branch, tag or commit, when --model is a Hugging Face repo id")
|
|
@@ -951,7 +1175,8 @@ def main(argv: list[str] | None = None) -> int:
|
|
| 951 |
p.add_argument("--dtype", choices=DTYPE_CHOICES, default="auto", help="what to compute in; auto (default): "
|
| 952 |
"bfloat16 on a GPU with native bfloat16, else float32 (bf16 weights upcast); float16 is not "
|
| 953 |
"supported")
|
| 954 |
-
p.add_argument("--batch-tokens", type=int,
|
|
|
|
| 955 |
p.add_argument("--max-chunks", type=int, default=0, help="score at most N evenly spaced chunks per document")
|
| 956 |
p.add_argument("--drop-line", help='"calibrated" (default), "default" (the schema\'s hard filters) or an '
|
| 957 |
"expression such as 'toxicity >= 4 or spam_seo >= 3'")
|
|
@@ -960,50 +1185,109 @@ def main(argv: list[str] | None = None) -> int:
|
|
| 960 |
p.add_argument("--no-chunks", action="store_true", help="leave out the per-chunk list of split documents")
|
| 961 |
p.add_argument("--group", type=int, default=256, help="documents scored together")
|
| 962 |
args = p.parse_args(argv)
|
| 963 |
-
|
|
|
|
| 964 |
p.error(f"--input {args.input}: no such file")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 965 |
|
| 966 |
t0 = time.time()
|
| 967 |
-
|
| 968 |
-
|
| 969 |
-
|
|
|
|
|
|
|
|
|
|
| 970 |
compute = str(next(model.parameters()).dtype).replace("torch.", "")
|
| 971 |
print(f"source1: loaded {model.weights_file} ({model.weights_dtype or '?'} weights) on {model.device}, computing "
|
| 972 |
f"in {compute}, in {time.time() - t0:.1f} s; drop line: {model.drop_line.source}", file=sys.stderr)
|
| 973 |
-
out = open(args.output, "w", encoding="utf-8") if args.output else sys.stdout
|
| 974 |
-
done = 0
|
| 975 |
t0 = time.time()
|
| 976 |
|
| 977 |
def flush(group: list[dict]) -> None:
|
| 978 |
-
nonlocal done
|
| 979 |
for rec, res in zip(group, model.score_batch(group, args.batch_tokens, text_field=args.text_field,
|
| 980 |
max_chunks=args.max_chunks)):
|
|
|
|
| 981 |
if args.no_chunks:
|
| 982 |
res.pop("chunks", None)
|
| 983 |
if "id" in rec:
|
| 984 |
res = {"id": rec["id"], **res}
|
| 985 |
-
|
| 986 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 987 |
print(f"source1: {done:,} documents, {done / max(time.time() - t0, 1e-9):.1f}/s", file=sys.stderr)
|
| 988 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 989 |
try:
|
| 990 |
-
group: list[dict] = []
|
| 991 |
try:
|
| 992 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 993 |
group.append(rec)
|
| 994 |
if len(group) >= args.group:
|
| 995 |
flush(group)
|
| 996 |
group = []
|
| 997 |
-
except BadInput as e:
|
| 998 |
if group:
|
| 999 |
-
flush(group)
|
| 1000 |
-
|
| 1001 |
-
|
| 1002 |
-
|
| 1003 |
-
|
| 1004 |
-
|
| 1005 |
-
|
| 1006 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1007 |
return 0
|
| 1008 |
|
| 1009 |
|
|
|
|
| 17 |
|
| 18 |
python source1.py --input docs.jsonl --output scores.jsonl # one JSON object per line, "text" field
|
| 19 |
python source1.py --input page.txt # a .txt file is one document
|
| 20 |
+
python source1.py --input docs.jsonl.gz # gzip, bzip2 and xz files are decompressed
|
| 21 |
|
| 22 |
Output: one flat dict per document
|
| 23 |
----------------------------------
|
|
|
|
| 43 |
|
| 44 |
How a document is scored (the same steps the model was trained and evaluated with)
|
| 45 |
------------------------------------------------------------------------------------
|
| 46 |
+
1. ``clean_text``: line endings to "\\n", ASCII control characters other than tab and newline dropped, Unicode NFC,
|
| 47 |
+
trailing spaces dropped, at most two blank lines in a row.
|
| 48 |
2. ``split_text``: a document longer than 7,808 tokens is split into balanced chunks, each ending at the most
|
| 49 |
natural boundary near its ideal end (headings, then paragraphs, lines, sentences, spaces; definitions in code).
|
| 50 |
3. ``build_input``: each chunk gets a one-line header, a blank line, then the chunk text:
|
|
|
|
| 52 |
Every training input had a Source line, almost always "dataset record", so that is the default; code files had
|
| 53 |
"<language> source file" (pass ``code_language="Python"``). Title is added when given, Part when the document has
|
| 54 |
more than one chunk. ``url`` is accepted but not shown to the model unless ``show_url=True``: no training input had
|
| 55 |
+
one. On 1,261 held-out benchmark chunks (495 graded held-out chunks from the test split and 766 exam chunks),
|
| 56 |
dropping the whole header moved overall by 0.03 on average (at most 0.655), dropping a title by 0.05 on the chunks
|
| 57 |
that had one; adding a URL moved it by up to 0.7 and did not improve its rank agreement with the independent
|
| 58 |
graders of the model card's evaluation.
|
|
|
|
| 86 |
|
| 87 |
import argparse
|
| 88 |
import ast
|
| 89 |
+
import bz2
|
| 90 |
+
import gzip
|
| 91 |
import json
|
| 92 |
+
import lzma
|
| 93 |
import math
|
| 94 |
import operator
|
| 95 |
import re
|
| 96 |
import sys
|
| 97 |
import time
|
| 98 |
import unicodedata
|
| 99 |
+
import zlib
|
| 100 |
from bisect import bisect_left
|
| 101 |
from collections.abc import Iterable, Iterator, Mapping
|
| 102 |
from pathlib import Path
|
|
|
|
| 110 |
MAX_LENGTH = 8192 # tokens per model input, <bos> and <eos> included
|
| 111 |
CHUNK_TOKENS = 7808 # document tokens per chunk: 8,192 minus 384 kept for the header (as in training)
|
| 112 |
DEFAULT_SOURCE = "dataset record"
|
| 113 |
+
DEFAULT_BATCH_TOKENS = 65536 # padded tokens per forward pass on a GPU
|
| 114 |
+
DEFAULT_BATCH_TOKENS_CPU = 16384 # on a CPU: a quarter of the memory, and no slower there
|
| 115 |
LEVELS = (0, 1, 2, 3, 4, 5)
|
| 116 |
# The backbone weights of each precision: bfloat16 (the default) and the full-precision float32 copy. The fp32 file
|
| 117 |
# follows transformers' variant naming (model.<variant>.safetensors), so AutoModel loads it with variant="fp32".
|
|
|
|
| 132 |
|
| 133 |
|
| 134 |
_JUNK = re.compile("[\x00-\x08\x0b\x0e-\x1f\x7f\ud800-\udfff\ufeff\u200b\ufffe\uffff]")
|
| 135 |
+
# The lookbehind keeps this linear: without it, a long run of spaces that ends in no newline is scanned again from
|
| 136 |
+
# each of its positions, so a page of spaces takes quadratic time. The result is the same.
|
| 137 |
+
_TRAILING_WS = re.compile(r"(?<![ \t])[ \t]+\n")
|
| 138 |
_BLANK_LINES = re.compile(r"\n{4,}")
|
| 139 |
_SPACE_RUN = re.compile(r"[ \t]{2,}")
|
| 140 |
_BLANK_RUN = re.compile(r"\n(?:[ \t]*\n){3,}")
|
|
|
|
| 142 |
|
| 143 |
def clean_text(text: str) -> str:
|
| 144 |
"""The document as the scorer sees it before chunking: "\\n" line endings (a form feed counts as a paragraph
|
| 145 |
+
break), ASCII control characters other than tab and newline, lone surrogates, zero-width spaces and byte-order
|
| 146 |
+
marks dropped (C1 control characters such as U+0085 are kept), Unicode NFC, no trailing spaces, at most two
|
| 147 |
+
blank lines in a row, no blank lines at either end."""
|
| 148 |
text = text.replace("\r\n", "\n").replace("\r", "\n").replace("\x0c", "\n\n")
|
| 149 |
text = _JUNK.sub("", text)
|
| 150 |
text = unicodedata.normalize("NFC", text)
|
|
|
|
| 305 |
ast.Gt: operator.gt, ast.GtE: operator.ge}
|
| 306 |
|
| 307 |
|
| 308 |
+
class DropLineError(ValueError):
|
| 309 |
+
"""A drop line that is not valid (see ``DropLine``)."""
|
| 310 |
+
|
| 311 |
+
|
| 312 |
class DropLine:
|
| 313 |
+
"""A drop line such as ``toxicity >= 4 or spam_seo >= 3.5 or boilerplate >= 4.5``: comparisons of a score (or
|
| 314 |
+
``overall``, ``tokens``, ``parts``) with a number or another score, and of a label with one of its values
|
| 315 |
+
(``format == 'news'``; ``==`` and ``!=`` only), joined with ``and`` / ``or`` / ``not`` and parentheses; ``True``
|
| 316 |
+
and ``False`` match always and never. A comparison with a missing score (None) is false, ``!=`` included. A
|
| 317 |
+
chunk or document is kept when the line does not match.
|
| 318 |
+
|
| 319 |
+
``names`` are the fields the line may use and ``labels`` maps each label field to its values; the other names
|
| 320 |
+
are numbers. The line is checked when it is created: a bare field name or constant used as a condition, a string
|
| 321 |
+
that is not a value of its label, a score compared with a string or a label with a number raise DropLineError
|
| 322 |
+
(a ValueError)."""
|
| 323 |
|
| 324 |
_NODES = (ast.Expression, ast.BoolOp, ast.And, ast.Or, ast.UnaryOp, ast.Not, ast.USub, ast.Compare, ast.Name,
|
| 325 |
ast.Load, ast.Constant, *_CMP)
|
| 326 |
|
| 327 |
+
def __init__(self, source: str, names: Iterable[str], labels: Mapping[str, Iterable[str]] | None = None):
|
| 328 |
self.source = source.strip()
|
| 329 |
+
self.names = set(names)
|
| 330 |
+
self.labels = {name: list(values) for name, values in (labels or {}).items()}
|
| 331 |
+
try:
|
| 332 |
+
tree = ast.parse(self.source, mode="eval")
|
| 333 |
+
except (SyntaxError, ValueError, RecursionError, MemoryError) as e:
|
| 334 |
+
raise DropLineError(f"drop line {source!r} is not a valid expression ({e})") from None
|
| 335 |
for node in ast.walk(tree):
|
| 336 |
if not isinstance(node, self._NODES):
|
| 337 |
+
raise DropLineError(f"{type(node).__name__} is not allowed in a drop line: {source!r}")
|
| 338 |
+
if isinstance(node, ast.Name) and node.id not in self.names:
|
| 339 |
+
raise DropLineError(f"unknown name {node.id!r} in drop line {source!r}")
|
| 340 |
self.tree = tree.body
|
| 341 |
+
try:
|
| 342 |
+
self._condition(self.tree)
|
| 343 |
+
except RecursionError:
|
| 344 |
+
raise DropLineError(f"drop line {source!r} is nested too deeply") from None
|
| 345 |
top_or = isinstance(self.tree, ast.BoolOp) and isinstance(self.tree.op, ast.Or)
|
| 346 |
self.terms = list(self.tree.values) if top_or else [self.tree]
|
| 347 |
+
# Evaluated once on stand-in scores, and once with every score missing, so that a line that cannot be
|
| 348 |
+
# evaluated fails here and not in the middle of a run.
|
| 349 |
+
try:
|
| 350 |
+
self.reasons({n: self.labels[n][0] if self.labels.get(n) else 0.0 for n in self.names})
|
| 351 |
+
self.reasons({})
|
| 352 |
+
except RecursionError:
|
| 353 |
+
raise DropLineError(f"drop line {source!r} is nested too deeply") from None
|
| 354 |
+
|
| 355 |
+
def _fail(self, node: ast.AST, problem: str) -> None:
|
| 356 |
+
raise DropLineError(f"{ast.unparse(node)!r} {problem}, in drop line {self.source!r}")
|
| 357 |
+
|
| 358 |
+
def _condition(self, node: ast.AST) -> None:
|
| 359 |
+
"""Check that ``node`` is a condition: a comparison, True / False, or and / or / not of conditions."""
|
| 360 |
+
if isinstance(node, ast.BoolOp):
|
| 361 |
+
for value in node.values:
|
| 362 |
+
self._condition(value)
|
| 363 |
+
elif isinstance(node, ast.UnaryOp) and isinstance(node.op, ast.Not):
|
| 364 |
+
self._condition(node.operand)
|
| 365 |
+
elif isinstance(node, ast.Compare):
|
| 366 |
+
kinds = [self._operand(x) for x in (node.left, *node.comparators)]
|
| 367 |
+
for op, (a, b) in zip(node.ops, zip(kinds, kinds[1:])):
|
| 368 |
+
self._check_pair(node, op, a, b)
|
| 369 |
+
elif not (isinstance(node, ast.Constant) and isinstance(node.value, bool)):
|
| 370 |
+
self._fail(node, "is not a condition: compare it with something, as in 'toxicity >= 4'")
|
| 371 |
+
|
| 372 |
+
def _operand(self, node: ast.AST) -> tuple[str, Any]:
|
| 373 |
+
"""What one side of a comparison is: ("label", name), ("str", value), ("num", name) for a numeric field or
|
| 374 |
+
("num", None) for a number."""
|
| 375 |
+
if isinstance(node, ast.Name):
|
| 376 |
+
return ("label", node.id) if node.id in self.labels else ("num", node.id)
|
| 377 |
+
if isinstance(node, ast.UnaryOp) and isinstance(node.op, ast.USub):
|
| 378 |
+
kind = self._operand(node.operand)
|
| 379 |
+
if kind[0] != "num":
|
| 380 |
+
self._fail(node, "negates something that is not a number")
|
| 381 |
+
return kind
|
| 382 |
+
if isinstance(node, ast.Constant):
|
| 383 |
+
v = node.value
|
| 384 |
+
if isinstance(v, str):
|
| 385 |
+
return ("str", v)
|
| 386 |
+
if (isinstance(v, int) and not isinstance(v, bool)) or (isinstance(v, float) and math.isfinite(v)):
|
| 387 |
+
return ("num", None)
|
| 388 |
+
self._fail(node, "is not allowed in a comparison: use a finite number or a label value in quotes")
|
| 389 |
+
self._fail(node, "is not allowed in a comparison: compare fields, numbers and label values")
|
| 390 |
+
raise AssertionError # not reached
|
| 391 |
+
|
| 392 |
+
def _check_pair(self, node: ast.Compare, op: ast.cmpop, a: tuple[str, Any], b: tuple[str, Any]) -> None:
|
| 393 |
+
kinds = {a[0], b[0]}
|
| 394 |
+
if all(k == "str" or name is None for k, name in (a, b)):
|
| 395 |
+
self._fail(node, "compares two constants")
|
| 396 |
+
if kinds == {"num"}:
|
| 397 |
+
return
|
| 398 |
+
if kinds == {"label", "str"}:
|
| 399 |
+
(_, label), (_, value) = (a, b) if a[0] == "label" else (b, a)
|
| 400 |
+
if not isinstance(op, (ast.Eq, ast.NotEq)):
|
| 401 |
+
self._fail(node, f"compares the label {label} by order: use == or !=")
|
| 402 |
+
if value not in self.labels[label]:
|
| 403 |
+
self._fail(node, f"uses {value!r}, which is not a value of {label} (one of: "
|
| 404 |
+
f"{', '.join(self.labels[label])})")
|
| 405 |
+
return
|
| 406 |
+
if kinds == {"label"}:
|
| 407 |
+
self._fail(node, "compares two labels: compare a label with one of its values, as in format == 'news'")
|
| 408 |
+
self._fail(node, "compares a score with a string, or a label with a number")
|
| 409 |
|
| 410 |
def reasons(self, env: dict) -> list[str]:
|
| 411 |
"""The terms of the line that match ``env`` (each ``or`` branch on its own); [] means keep."""
|
|
|
|
| 432 |
left = self._eval(node.left, env)
|
| 433 |
for op, comp in zip(node.ops, node.comparators):
|
| 434 |
right = self._eval(comp, env)
|
| 435 |
+
if left is None or right is None or not _CMP[type(op)](left, right):
|
| 436 |
+
return False # a comparison with a missing value is false, != included
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 437 |
left = right
|
| 438 |
return True
|
| 439 |
+
raise DropLineError(f"unsupported drop-line node {type(node).__name__}")
|
| 440 |
|
| 441 |
|
| 442 |
def _r3(x: float | None) -> float | None:
|
|
|
|
| 653 |
return calibration
|
| 654 |
|
| 655 |
|
| 656 |
+
def _make_drop_line(config: dict, calibration: dict | None, drop_line: str | None = None) -> DropLine:
|
| 657 |
+
"""The ``DropLine`` of ``drop_line`` for the schema of source1.json (``config``): None or "calibrated" =
|
| 658 |
+
calibration.json's line, "default" = the schema's hard filters, anything else = the expression itself, checked
|
| 659 |
+
against the schema's fields and label values (ValueError when it is not a valid drop line)."""
|
| 660 |
+
schema = config["schema"]
|
| 661 |
+
if drop_line in (None, "calibrated"):
|
| 662 |
+
drop_line = ((calibration or {}).get("drop_line") or {}).get("line")
|
| 663 |
+
if not drop_line:
|
| 664 |
+
raise ValueError("no calibrated drop line (calibration.json missing or incomplete); pass "
|
| 665 |
+
"drop_line='default' or your own drop line")
|
| 666 |
+
elif drop_line == "default":
|
| 667 |
+
drop_line = " or ".join((schema.get("composite") or {}).get("hard_filters") or []) or "False"
|
| 668 |
+
labels = {name: list(spec["values"]) for name, spec in schema["labels"].items()}
|
| 669 |
+
numbers = [n for g in ("quality", "red_flags", "gated") for n in schema.get(g, {})] + ["overall", "tokens", "parts"]
|
| 670 |
+
return DropLine(drop_line, [*labels, *numbers], labels)
|
| 671 |
+
|
| 672 |
+
|
| 673 |
class Source1(nn.Module):
|
| 674 |
"""Source-1: the mmBERT-base encoder, mean pooling and one linear head per field. Build it with
|
| 675 |
``Source1.from_pretrained``; score documents with ``score`` / ``score_batch``."""
|
|
|
|
| 693 |
self.labels = {name: list(spec["values"]) for name, spec in self.schema["labels"].items()}
|
| 694 |
self.fields = list(self.labels) + [n for g in ("quality", "red_flags", "gated") for n in self.schema.get(g, {})]
|
| 695 |
self.calibration = calibration or {}
|
| 696 |
+
self.drop_line = _make_drop_line(config, self.calibration, drop_line)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 697 |
self.offsets = {k: float(v["offset"]) for k, v in (self.calibration.get("offsets") or {}).items()}
|
| 698 |
self.apply_offsets = apply_offsets
|
| 699 |
self.show_url = show_url
|
|
|
|
| 753 |
raise FileNotFoundError(f"{path} lacks {', '.join(missing)}")
|
| 754 |
config = json.loads((path / "source1.json").read_text(encoding="utf-8"))
|
| 755 |
calibration = _load_calibration(path, drop_line, apply_offsets)
|
| 756 |
+
_make_drop_line(config, calibration, drop_line) # a bad drop line fails here, before the weights load
|
| 757 |
tok = Tokenizer.from_file(str(path / "tokenizer.json"))
|
| 758 |
tok.no_truncation()
|
| 759 |
tok.no_padding()
|
| 760 |
variant = {"variant": "fp32"} if precision == "fp32" else {}
|
| 761 |
+
backbone_kwargs.pop("output_loading_info", None) # always requested, to check the weights
|
| 762 |
+
backbone, info = AutoModel.from_pretrained(str(path), **{_dtype_kwarg(): dtype}, **variant, **backbone_kwargs,
|
| 763 |
+
output_loading_info=True)
|
| 764 |
+
keys = ("missing_keys", "unexpected_keys", "mismatched_keys", "error_msgs")
|
| 765 |
+
bad = {k: info[k] for k in keys if info.get(k)}
|
| 766 |
+
if bad: # transformers would otherwise fill the missing weights with random values, without an error
|
| 767 |
+
detail = "; ".join(f"{len(v)} {k.replace('_', ' ')} ({', '.join(sorted(map(str, v))[:3])}"
|
| 768 |
+
f"{', ...' if len(v) > 3 else ''})" for k, v in bad.items())
|
| 769 |
+
raise RuntimeError(f"Source-1: the backbone weights in {path / weights} do not match the model: {detail}. "
|
| 770 |
+
"Download the weights again, and use the package versions in requirements.txt")
|
| 771 |
model = cls(backbone, tok, config, calibration, drop_line, apply_offsets, show_url)
|
| 772 |
model.heads.load_state_dict(load_file(str(path / "heads.safetensors")))
|
| 773 |
model.heads.to(dtype)
|
|
|
|
| 859 |
rec["drop_reasons"] = reasons
|
| 860 |
return rec
|
| 861 |
|
| 862 |
+
def default_batch_tokens(self) -> int:
|
| 863 |
+
"""The padded tokens per forward pass when ``batch_tokens`` is None: 16,384 on a CPU, 65,536 on a GPU."""
|
| 864 |
+
return DEFAULT_BATCH_TOKENS_CPU if self.device.type == "cpu" else DEFAULT_BATCH_TOKENS
|
| 865 |
+
|
| 866 |
+
def score_inputs(self, inputs: list[str], batch_tokens: int | None = None) -> list[dict]:
|
| 867 |
"""Score ready-made model inputs (``build_input`` output: header, blank line, chunk text), one dict per
|
| 868 |
input with the 13 fields, overall, keep, drop_reasons, input_tokens and truncated.
|
| 869 |
|
| 870 |
+
Inputs are sorted by length and batched with at most ``batch_tokens`` padded tokens per forward pass
|
| 871 |
+
(default: 16,384 on a CPU, 65,536 on a GPU); a batch that runs out of GPU memory is split in half and
|
| 872 |
+
retried."""
|
| 873 |
+
if batch_tokens is None:
|
| 874 |
+
batch_tokens = self.default_batch_tokens()
|
| 875 |
encoded = self.encode(inputs)
|
| 876 |
ids = [e[0] for e in encoded]
|
| 877 |
results: list[dict | None] = [None] * len(ids)
|
|
|
|
| 901 |
run(batch)
|
| 902 |
return results # type: ignore[return-value]
|
| 903 |
|
| 904 |
+
def score_batch(self, docs: Iterable[str | bytes | dict], batch_tokens: int | None = None, *,
|
| 905 |
text_field: str = "text", max_chunks: int = 0) -> list[dict]:
|
| 906 |
"""Score a list of documents: strings (bytes are decoded as UTF-8), or dicts with the text under
|
| 907 |
``text_field`` and optionally ``title``, ``url``, ``source_type`` and ``code_language`` (see ``score``).
|
| 908 |
A None text is scored as an empty document (keep False, drop_reasons ["empty text"]). Chunks of all
|
| 909 |
documents are batched together. ``max_chunks`` > 0 scores only that many evenly spaced chunks of a long
|
| 910 |
+
document (the Part numbers still count every chunk). ``batch_tokens``: see ``score_inputs``.
|
| 911 |
|
| 912 |
In bfloat16 a document's scores can shift slightly (up to about 0.04 on overall, 0.10 on a single field)
|
| 913 |
depending on which other documents share its batch, because the batch shape changes the kernels' rounding;
|
|
|
|
| 943 |
|
| 944 |
def score(self, text: str | bytes, title: str | None = None, url: str | None = None, *,
|
| 945 |
source_type: str | None = None, code_language: str | None = None, max_chunks: int = 0,
|
| 946 |
+
batch_tokens: int | None = None) -> dict:
|
| 947 |
"""Score one document (a str; bytes are decoded as UTF-8; anything else raises TypeError).
|
| 948 |
|
| 949 |
title: shown to the model in the header when given (as in training, where about a quarter of inputs had one).
|
|
|
|
| 981 |
|
| 982 |
|
| 983 |
class BadInput(ValueError):
|
| 984 |
+
"""A record of the input file that cannot be scored (the message starts with file:line), or, with
|
| 985 |
+
``skippable=False``, an input file that cannot be read at all."""
|
| 986 |
+
|
| 987 |
+
def __init__(self, message: str, skippable: bool = True):
|
| 988 |
+
super().__init__(message)
|
| 989 |
+
self.skippable = skippable
|
| 990 |
+
|
| 991 |
+
|
| 992 |
+
# Compressed inputs read transparently, recognized by their first bytes whatever their name.
|
| 993 |
+
_DECOMPRESS = {"gzip": gzip.open, "bzip2": bz2.open, "xz": lzma.open}
|
| 994 |
+
_COMPRESSED_SUFFIXES = (".gz", ".bz2", ".xz") # docs.jsonl.gz is read as .jsonl
|
| 995 |
+
_UNREADABLE = {".zst": "is zstd-compressed: decompress it first (zstd -d)",
|
| 996 |
+
".zstd": "is zstd-compressed: decompress it first (zstd -d)",
|
| 997 |
+
".parquet": "is a Parquet file: convert it to JSON Lines first",
|
| 998 |
+
".zip": "is a zip archive: extract it first",
|
| 999 |
+
".7z": "is a 7z archive: extract it first"}
|
| 1000 |
+
_READ_ERRORS = (OSError, EOFError, zlib.error, lzma.LZMAError) # what damaged or truncated compressed files raise
|
| 1001 |
+
_SNIFF_BYTES = 65536 # how much of an input file _input_problem looks at
|
| 1002 |
+
_MAX_INVALID = 0.2 # share of invalid UTF-8 sequences among the characters above which an input file is refused
|
| 1003 |
+
_MAX_NUL = 0.01 # share of NUL bytes above which an input file is refused as binary
|
| 1004 |
+
|
| 1005 |
+
|
| 1006 |
+
def _compression(head: bytes) -> str | None:
|
| 1007 |
+
"""The compression that a file's first bytes show: "gzip", "bzip2", "xz", "zstd" or None."""
|
| 1008 |
+
if head.startswith(b"\x1f\x8b"):
|
| 1009 |
+
return "gzip"
|
| 1010 |
+
if head[:3] == b"BZh" and b"1" <= head[3:4] <= b"9" and head[4:10] in (b"1AY&SY", b"\x17rE8P\x90"):
|
| 1011 |
+
return "bzip2"
|
| 1012 |
+
if head.startswith(b"\xfd7zXZ\x00"):
|
| 1013 |
+
return "xz"
|
| 1014 |
+
if head.startswith(b"\x28\xb5\x2f\xfd"):
|
| 1015 |
+
return "zstd"
|
| 1016 |
+
return None
|
| 1017 |
+
|
| 1018 |
+
|
| 1019 |
+
def _open_input(path: Path) -> Any:
|
| 1020 |
+
"""``path`` opened for reading bytes, decompressed when it is a gzip, bzip2 or xz file."""
|
| 1021 |
+
with open(path, "rb") as f:
|
| 1022 |
+
head = f.read(10)
|
| 1023 |
+
return _DECOMPRESS.get(_compression(head) or "", open)(path, "rb")
|
| 1024 |
+
|
| 1025 |
+
|
| 1026 |
+
def _format_suffix(path: Path) -> str:
|
| 1027 |
+
"""The suffix that says how to read ``path``: its last one, or the one before .gz, .bz2 or .xz."""
|
| 1028 |
+
suffixes = [s.lower() for s in path.suffixes]
|
| 1029 |
+
if suffixes and suffixes[-1] in _COMPRESSED_SUFFIXES:
|
| 1030 |
+
suffixes.pop()
|
| 1031 |
+
return suffixes[-1] if suffixes else ""
|
| 1032 |
+
|
| 1033 |
+
|
| 1034 |
+
def _input_problem(path: Path) -> str | None:
|
| 1035 |
+
"""Why ``path`` cannot be scored as UTF-8 text or JSON (zstd, Parquet, an archive, binary, UTF-16, mostly invalid
|
| 1036 |
+
UTF-8, damaged), judged from its name and its first 64 KiB once decompressed; None when it looks readable."""
|
| 1037 |
+
if path.suffix.lower() in _UNREADABLE:
|
| 1038 |
+
return _UNREADABLE[path.suffix.lower()]
|
| 1039 |
+
try:
|
| 1040 |
+
with _open_input(path) as f:
|
| 1041 |
+
head = f.read(_SNIFF_BYTES)
|
| 1042 |
+
except _READ_ERRORS as e:
|
| 1043 |
+
return f"cannot be read ({e})"
|
| 1044 |
+
if _compression(head) == "zstd":
|
| 1045 |
+
return _UNREADABLE[".zst"]
|
| 1046 |
+
if head.startswith((b"\xff\xfe", b"\xfe\xff")):
|
| 1047 |
+
return "is UTF-16 or UTF-32: save it as UTF-8"
|
| 1048 |
+
if head.count(b"\x00") > _MAX_NUL * len(head): # a stray NUL in a text is dropped like other control characters
|
| 1049 |
+
return "holds NUL bytes, so it is not UTF-8 text (binary, UTF-16, or compressed in a format not read here)"
|
| 1050 |
+
text = head.decode("utf-8", errors="replace")
|
| 1051 |
+
invalid = text.count("\ufffd") - head.count("\ufffd".encode()) # a U+FFFD already in the text is valid UTF-8
|
| 1052 |
+
if invalid >= 8 and invalid > _MAX_INVALID * len(text):
|
| 1053 |
+
return f"is not UTF-8 text: {invalid / len(text):.0%} of its first {len(text):,} characters are invalid UTF-8"
|
| 1054 |
+
return None
|
| 1055 |
+
|
| 1056 |
+
|
| 1057 |
+
def _json_error(e: BaseException) -> str:
|
| 1058 |
+
"""Why json.loads failed, in a few words."""
|
| 1059 |
+
if isinstance(e, json.JSONDecodeError):
|
| 1060 |
+
return f"{e.msg} at column {e.colno}"
|
| 1061 |
+
if isinstance(e, RecursionError):
|
| 1062 |
+
return "nested too deeply"
|
| 1063 |
+
return str(e).split(";")[0] # e.g. an integer of more than 4,300 digits
|
| 1064 |
+
|
| 1065 |
+
|
| 1066 |
+
def _unwritable(value: Any) -> str | None:
|
| 1067 |
+
"""Why ``value`` cannot be written to the JSON Lines output, or None when it can."""
|
| 1068 |
+
try:
|
| 1069 |
+
json.dumps(value, ensure_ascii=False, allow_nan=False).encode("utf-8")
|
| 1070 |
+
except UnicodeEncodeError:
|
| 1071 |
+
return "holds a lone surrogate, which UTF-8 cannot encode"
|
| 1072 |
+
except ValueError:
|
| 1073 |
+
return "holds NaN or an infinite number, which JSON cannot hold"
|
| 1074 |
+
except RecursionError:
|
| 1075 |
+
return "is nested too deeply"
|
| 1076 |
+
return None
|
| 1077 |
|
| 1078 |
|
| 1079 |
def _read_docs(path: Path, text_field: str, skip_bad: bool = False) -> Iterator[dict]:
|
| 1080 |
"""Documents of an input file: .jsonl / .ndjson (one JSON object per line), .json (a JSON array of objects, one
|
| 1081 |
+
object, or JSON Lines), or any other file as one plain-text document; gzip, bzip2 and xz files are decompressed
|
| 1082 |
+
(docs.jsonl.gz is read as .jsonl). Invalid UTF-8 is replaced, with a warning. A bad record (not a JSON object,
|
| 1083 |
+
no string text, an id that cannot be written back) raises BadInput, or with ``skip_bad`` is reported on stderr
|
| 1084 |
+
and skipped. A file that cannot be read at all (see ``_input_problem``) raises BadInput with skippable=False.
|
| 1085 |
+
A null text is kept and scored as an empty document."""
|
| 1086 |
|
| 1087 |
def usable(rec: Any, where: str) -> bool:
|
| 1088 |
if not isinstance(rec, dict):
|
|
|
|
| 1091 |
problem = f"no {text_field!r} field"
|
| 1092 |
elif rec[text_field] is not None and not isinstance(rec[text_field], str):
|
| 1093 |
problem = f"{text_field!r} is a {type(rec[text_field]).__name__}, not a string"
|
| 1094 |
+
elif "id" in rec and (why := _unwritable(rec["id"])):
|
| 1095 |
+
problem = f"the id {why}"
|
| 1096 |
else:
|
| 1097 |
return True
|
| 1098 |
if not skip_bad:
|
|
|
|
| 1108 |
file=sys.stderr)
|
| 1109 |
return raw.decode("utf-8", errors="replace")
|
| 1110 |
|
| 1111 |
+
def read_all() -> bytes:
|
| 1112 |
+
try:
|
| 1113 |
+
with _open_input(path) as f:
|
| 1114 |
+
return f.read()
|
| 1115 |
+
except _READ_ERRORS as e:
|
| 1116 |
+
raise BadInput(f"{path} cannot be read ({e})", skippable=False) from None
|
| 1117 |
+
|
| 1118 |
+
problem = _input_problem(path)
|
| 1119 |
+
if problem:
|
| 1120 |
+
raise BadInput(f"{path} {problem}", skippable=False)
|
| 1121 |
+
suffix = _format_suffix(path)
|
| 1122 |
if suffix not in (".jsonl", ".ndjson", ".json"):
|
| 1123 |
+
yield {text_field: decode(read_all(), str(path)), "id": path.name}
|
| 1124 |
return
|
| 1125 |
if suffix == ".json":
|
| 1126 |
try:
|
| 1127 |
+
data = json.loads(decode(read_all(), str(path)).lstrip(""))
|
| 1128 |
+
except (ValueError, RecursionError): # JSONDecodeError is a ValueError
|
| 1129 |
+
data = None # not one JSON value (or too large or too deep to read as one): read it as JSON Lines below
|
| 1130 |
if data is not None:
|
| 1131 |
for n, rec in enumerate(data if isinstance(data, list) else [data]):
|
| 1132 |
if usable(rec, f"{path}[{n}]"):
|
| 1133 |
yield rec
|
| 1134 |
return
|
| 1135 |
+
n = 0
|
| 1136 |
+
try:
|
| 1137 |
+
with _open_input(path) as f:
|
| 1138 |
+
for n, raw in enumerate(f, 1):
|
| 1139 |
+
where = f"{path}:{n}"
|
| 1140 |
+
line = decode(raw, where)
|
| 1141 |
+
if n == 1:
|
| 1142 |
+
line = line.lstrip("")
|
| 1143 |
+
if not line.strip():
|
| 1144 |
+
continue
|
| 1145 |
+
try:
|
| 1146 |
+
rec = json.loads(line)
|
| 1147 |
+
except (ValueError, RecursionError) as e: # also integers too long to read, and very deep nesting
|
| 1148 |
+
problem = f"invalid JSON ({_json_error(e)})"
|
| 1149 |
+
if not skip_bad:
|
| 1150 |
+
raise BadInput(f"{where}: {problem}") from None
|
| 1151 |
+
print(f"source1: skipped {where}: {problem}", file=sys.stderr)
|
| 1152 |
+
continue
|
| 1153 |
+
if usable(rec, where):
|
| 1154 |
+
yield rec
|
| 1155 |
+
except _READ_ERRORS as e:
|
| 1156 |
+
raise BadInput(f"{path} cannot be read after line {n} ({e})", skippable=False) from None
|
| 1157 |
|
| 1158 |
|
| 1159 |
def main(argv: list[str] | None = None) -> int:
|
|
|
|
| 1161 |
p.add_argument("--model", default=str(Path(__file__).resolve().parent),
|
| 1162 |
help="Source-1 directory or Hugging Face repo id (default: this file's directory)")
|
| 1163 |
p.add_argument("--input", required=True, help=".jsonl (one document per line), .json (an array of objects) or "
|
| 1164 |
+
"a text file (one document); gzip, bzip2 and xz files are decompressed (docs.jsonl.gz)")
|
| 1165 |
p.add_argument("--text-field", default="text", help="JSON field holding the text (default: text); "
|
| 1166 |
"title, url, source_type and code_language fields are used when present")
|
| 1167 |
+
p.add_argument("--output", help="output .jsonl (default: standard output); written to <output>.tmp and renamed "
|
| 1168 |
+
"at the end, so a run that fails leaves an older output as it was")
|
| 1169 |
p.add_argument("--skip-bad", action="store_true", help="skip (and report on stderr) records that are not valid "
|
| 1170 |
"JSON objects with a string text, instead of stopping")
|
| 1171 |
p.add_argument("--revision", help="branch, tag or commit, when --model is a Hugging Face repo id")
|
|
|
|
| 1175 |
p.add_argument("--dtype", choices=DTYPE_CHOICES, default="auto", help="what to compute in; auto (default): "
|
| 1176 |
"bfloat16 on a GPU with native bfloat16, else float32 (bf16 weights upcast); float16 is not "
|
| 1177 |
"supported")
|
| 1178 |
+
p.add_argument("--batch-tokens", type=int, help="padded tokens per forward pass (default: "
|
| 1179 |
+
f"{DEFAULT_BATCH_TOKENS_CPU} on a CPU, {DEFAULT_BATCH_TOKENS} on a GPU)")
|
| 1180 |
p.add_argument("--max-chunks", type=int, default=0, help="score at most N evenly spaced chunks per document")
|
| 1181 |
p.add_argument("--drop-line", help='"calibrated" (default), "default" (the schema\'s hard filters) or an '
|
| 1182 |
"expression such as 'toxicity >= 4 or spam_seo >= 3'")
|
|
|
|
| 1185 |
p.add_argument("--no-chunks", action="store_true", help="leave out the per-chunk list of split documents")
|
| 1186 |
p.add_argument("--group", type=int, default=256, help="documents scored together")
|
| 1187 |
args = p.parse_args(argv)
|
| 1188 |
+
source = Path(args.input)
|
| 1189 |
+
if not source.is_file():
|
| 1190 |
p.error(f"--input {args.input}: no such file")
|
| 1191 |
+
problem = _input_problem(source)
|
| 1192 |
+
if problem:
|
| 1193 |
+
p.error(f"--input {args.input} {problem}")
|
| 1194 |
+
if args.batch_tokens is not None and args.batch_tokens < 1:
|
| 1195 |
+
p.error("--batch-tokens must be at least 1")
|
| 1196 |
+
# Checked before the model loads. A regular output file is written as <output>.tmp, which replaces the output
|
| 1197 |
+
# only when the run succeeds; a device or pipe (such as /dev/stdout) is written directly.
|
| 1198 |
+
target = tmp = None
|
| 1199 |
+
if args.output:
|
| 1200 |
+
out_path = Path(args.output)
|
| 1201 |
+
if out_path.is_dir():
|
| 1202 |
+
p.error(f"--output {args.output} is a folder; give a file name")
|
| 1203 |
+
if out_path.exists() and out_path.samefile(source):
|
| 1204 |
+
p.error("--output is the same file as --input; refusing to overwrite it")
|
| 1205 |
+
if not out_path.resolve().parent.is_dir():
|
| 1206 |
+
p.error(f"--output {args.output}: the folder does not exist")
|
| 1207 |
+
if out_path.is_file() or not out_path.exists():
|
| 1208 |
+
target = out_path.resolve()
|
| 1209 |
+
tmp = target.with_name(target.name + ".tmp")
|
| 1210 |
+
if tmp.is_dir() or (tmp.exists() and tmp.samefile(source)):
|
| 1211 |
+
p.error(f"--output {args.output}: the temporary file it is written to first, {tmp}, is "
|
| 1212 |
+
+ ("a folder" if tmp.is_dir() else "the --input file"))
|
| 1213 |
|
| 1214 |
t0 = time.time()
|
| 1215 |
+
try:
|
| 1216 |
+
model = Source1.from_pretrained(args.model, device=args.device, dtype=args.dtype, precision=args.precision,
|
| 1217 |
+
drop_line=args.drop_line, apply_offsets=args.apply_offsets,
|
| 1218 |
+
show_url=args.show_url, revision=args.revision)
|
| 1219 |
+
except DropLineError as e:
|
| 1220 |
+
p.error(str(e))
|
| 1221 |
compute = str(next(model.parameters()).dtype).replace("torch.", "")
|
| 1222 |
print(f"source1: loaded {model.weights_file} ({model.weights_dtype or '?'} weights) on {model.device}, computing "
|
| 1223 |
f"in {compute}, in {time.time() - t0:.1f} s; drop line: {model.drop_line.source}", file=sys.stderr)
|
| 1224 |
+
out = open(tmp or args.output, "w", encoding="utf-8") if args.output else sys.stdout
|
| 1225 |
+
done = seen = 0
|
| 1226 |
t0 = time.time()
|
| 1227 |
|
| 1228 |
def flush(group: list[dict]) -> None:
|
| 1229 |
+
nonlocal done, seen
|
| 1230 |
for rec, res in zip(group, model.score_batch(group, args.batch_tokens, text_field=args.text_field,
|
| 1231 |
max_chunks=args.max_chunks)):
|
| 1232 |
+
seen += 1
|
| 1233 |
if args.no_chunks:
|
| 1234 |
res.pop("chunks", None)
|
| 1235 |
if "id" in rec:
|
| 1236 |
res = {"id": rec["id"], **res}
|
| 1237 |
+
try: # the reader lets no unwritable id through; this keeps a half-written line out of the output
|
| 1238 |
+
line = json.dumps(res, ensure_ascii=False, allow_nan=False)
|
| 1239 |
+
line.encode("utf-8")
|
| 1240 |
+
except (ValueError, RecursionError) as e:
|
| 1241 |
+
where = f"document {seen:,}" + (f" (id {rec['id']!r})" if "id" in rec else "")
|
| 1242 |
+
if not args.skip_bad:
|
| 1243 |
+
raise BadInput(f"{where}: its scores cannot be written as JSON ({e})") from None
|
| 1244 |
+
print(f"source1: skipped {where}: its scores cannot be written as JSON ({e})", file=sys.stderr)
|
| 1245 |
+
continue
|
| 1246 |
+
out.write(line + "\n")
|
| 1247 |
+
done += 1
|
| 1248 |
print(f"source1: {done:,} documents, {done / max(time.time() - t0, 1e-9):.1f}/s", file=sys.stderr)
|
| 1249 |
|
| 1250 |
+
def written() -> str:
|
| 1251 |
+
"""What a run that stopped early left behind."""
|
| 1252 |
+
what = f"The {done:,} documents before it were" if done != 1 else "The document before it was"
|
| 1253 |
+
if tmp is None:
|
| 1254 |
+
return f"{what} written"
|
| 1255 |
+
if not done:
|
| 1256 |
+
tmp.unlink(missing_ok=True)
|
| 1257 |
+
return f"Nothing was written, and {args.output} was not changed"
|
| 1258 |
+
return f"{what} written to {tmp}, and {args.output} was not changed"
|
| 1259 |
+
|
| 1260 |
+
docs = _read_docs(source, args.text_field, args.skip_bad)
|
| 1261 |
+
group: list[dict] = []
|
| 1262 |
try:
|
|
|
|
| 1263 |
try:
|
| 1264 |
+
while True:
|
| 1265 |
+
try:
|
| 1266 |
+
rec = next(docs)
|
| 1267 |
+
except StopIteration:
|
| 1268 |
+
break
|
| 1269 |
+
except Exception:
|
| 1270 |
+
if group:
|
| 1271 |
+
flush(group) # the documents read before a bad record or a read error are still scored
|
| 1272 |
+
raise
|
| 1273 |
group.append(rec)
|
| 1274 |
if len(group) >= args.group:
|
| 1275 |
flush(group)
|
| 1276 |
group = []
|
|
|
|
| 1277 |
if group:
|
| 1278 |
+
flush(group)
|
| 1279 |
+
finally:
|
| 1280 |
+
if out is not sys.stdout:
|
| 1281 |
+
out.close()
|
| 1282 |
+
except BadInput as e:
|
| 1283 |
+
raise SystemExit(f"source1: {e}. {written()}"
|
| 1284 |
+
+ ("; --skip-bad skips bad records" if e.skippable else "")) from None
|
| 1285 |
+
except BaseException as e:
|
| 1286 |
+
if tmp is not None:
|
| 1287 |
+
print(f"source1: stopped by {type(e).__name__}. {written()}", file=sys.stderr)
|
| 1288 |
+
raise
|
| 1289 |
+
if tmp is not None:
|
| 1290 |
+
tmp.replace(target)
|
| 1291 |
return 0
|
| 1292 |
|
| 1293 |
|