Codebook
Column definitions for the public release dataset. 77 columns across 14 groups.
Provenance
Every column is tagged with one of four provenance flags so you can tell at a glance how the value was produced. Hover the badge on any row for the full derivation method.
collected · 8 Direct from publisher HTML
external · 39 From authoritative API (ROR, OpenAlex, DOAJ, COPE, NPI)
inferred · 7 Algorithmic / ML inference — may carry bias
computed · 23 Derived arithmetically from other columns
Inferred columns (especially gender) carry systematic bias against non-Latin-script names — see the methodology page for coverage details.
Core (11 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| publisher | Publisher | string | collected | Publisher name (e.g., Elsevier, Springer Nature, Wiley), known from the collection configuration for each source. |
| journal | Journal | string | collected | Journal title as listed on the publisher website. |
| editor | Name | string | collected | Editor full name as it appeared on the editorial board page, after mojibake repair and light cleaning. |
| first_name | First name | string | computed | Given name extracted from `editor` via the probablepeople ML name parser. Used as the lookup key for WGND 2.0 gender inference; preserved in the public release so users can re-use the parsed value without rerunning probablepeople. |
| last_name | Last name | string | computed | Surname extracted from `editor` via the probablepeople ML name parser. Empty when probablepeople could not parse the string; the original `editor` field is always preserved. |
| role | Role (raw) | string | collected | Role label as listed on the publisher page, unmodified. |
| role_std | Role (standardized) | string | computed | Standardized role, one of exactly ten canonical values: editor_in_chief, senior_editor, deputy_editor, associate_editor, section_editor, reviewing_editor, guest_editor, editorial_board_member, editor, other. The mapping is deterministic from the raw role string. 'senior_editor' was introduced in v3.0.0; 'other' absorbs roughly 6% of positions that do not match the canonical set. |
| affiliation | Affiliation | string | collected | Institutional affiliation as listed, after universal cleaning (strips roles, credentials, dates, junk). Not canonicalised; see ror_name for the authoritative form. |
| orcid | ORCID | string | collected | 16-digit ORCID iD. Preferred from the collected value when available; otherwise backfilled from OpenAlex or the ORCID API (see orcid_source). |
| source_url | Source URL | string | collected | Public publisher URL where this record was collected. |
| scraped_at | Collected at | datetime | collected | ISO-8601 timestamp of when this record was collected. The column name is retained for backward compatibility with earlier releases. |
Identity (1 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| editor_key | Editor key | string | computed | Composite identity key used to deduplicate editors across boards; the basis for the 744,842 unique-editor count. Resolved in priority order and prefixed with the tier that produced it: 'orcid:<orcid>|<ror>' when an ORCID iD is available (346,713 records, 250,217 unique editors); 'ror:<ror_id>|<name>' when there is no iD but the affiliation resolved to ROR (293,742 records, 235,869 unique); 'aff:<affiliation>|<name>' otherwise (281,642 records, 258,756 unique, of which 51,695 have an empty affiliation and so fall back to name alone). Published in v3.0.0 at reviewer request so the unique-editor count is reproducible without reconstructing the key. Because the lowest tier is name-based, distinct people who share a name and affiliation string can still collide there; the ORCID and ROR tiers are the reliable ones. |
Gender (5 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| gender | Gender | string | inferred | Country-aware inference from the editor's first name and ror_country using WGND 2.0 (Raffo & Lax-Martinez, WIPO 2021; Harvard Dataverse DOI 10.7910/DVN/MSEGSJ, roughly 3.5 million names across 195 countries). The same name can resolve differently by country (for example 'Andrea' is typically male in Italy and female in the United States). gender-guesser is a tertiary fallback for names absent from WGND. Values: male, female, unknown. The legacy 'andy' (androgynous) value was retired in v3.0.0. Self-reported gender is NOT available in this dataset; every value is inferred and should be treated as such. |
| gender_raw | Gender (raw) | string | inferred | Backwards-compatibility column kept since v1; mirrors gender exactly (WGND has no mostly_male / mostly_female granularity). Retained so v1-era scripts keep running; new work should read gender. |
| gender_prob | Gender weight | float | inferred | WGND weight in [0, 1]: the probability that an individual with this first name in this country has the inferred gender. Replaces the legacy 3-bucket {0.0, 0.75, 1.0} confidence; downstream filters should use thresholds like `>= 0.95`. For gender-guesser fallback rows, 1.0 = certain, 0.75 = mostly_*, 0.0 = unknown. |
| gender_nobs | WGND sample size | int | inferred | Number of observed individuals in the WGND cell (name × country) underlying the inference. Useful for filtering low-confidence cells in analysis. 0 for gender-guesser-fallback rows. |
| gender_source | Gender source | string | inferred | Provenance of the inference: `wgnd_country` (matched on first_name + ror_country, strongest signal), `wgnd_global` (matched on first_name only, used when country was missing or the country-specific cell was empty), `gender_guesser` (tertiary fallback, country-blind), or `unknown`. |
Institution (ROR) (8 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| ror_id | ROR ID | string | external | Research Organization Registry identifier resolved from the raw affiliation string via the ROR /organizations?affiliation= fuzzy-match endpoint, subject to score and token-overlap guards (see CHANGELOG v2.0.0) that reject low-confidence matches. Seven organisations drew clusters of false matches (issue #30). v4.0.0 re-applied the v2.0.0 repair, which later re-enrichments had undone: a match to one of the seven is kept only when the affiliation names that organisation. |
| ror_name | Institution | string | external | Canonical institution name from the ROR v2 record (ror_display name). |
| ror_country | Country | string | external | Country name from ROR/GeoNames for the primary location of the institution. |
| ror_city | City | string | external | City of the primary location in the ROR record of the row's ror_id (GeoNames). Taken from ROR for every row with a ror_id since v4.0.1, which filled 5,537 rows left empty in v4.0.0. |
| ror_state | State/Province | string | external | State or province of the primary location in the ROR record of the row's ror_id, as ROR gives it: usually a name, occasionally a GeoNames code (e.g. SG.01); empty when the ROR record gives none. Taken from ROR for every row with a ror_id since v4.0.1. |
| org_type | Org type | string | external | Organization type from ROR: education, healthcare, government, facility, nonprofit, company, archive, or other. Taken from the first entry in the ROR record's types array. |
| latitude | Latitude | float | external | Geographic latitude of the institution from ROR/GeoNames. |
| longitude | Longitude | float | external | Geographic longitude of the institution from ROR/GeoNames. |
Classification (OpenAlex) (4 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| scientific_domain | Domain | string | external | OpenAlex domain of scientific_field (Health Sciences, Life Sciences, Physical Sciences or Social Sciences), or Multidisciplinary. Journal-level: identical for every editor of a journal. |
| scientific_field | Field | string | external | Journal-level OpenAlex field (e.g. Medicine, Engineering, Psychology), identical for every editor of a journal: it describes the journal, not the editor's own research. Since v4.0.0 it is the field of most of the journal's own research articles with an abstract, by primary topic (falling back to all articles, then all works, below 30 classified works), provided that field covers at least 40% of those articles or the journal's earlier topic was one of OpenAlex's catch-all topics. Otherwise the journal keeps its v3.0.0 classification, which came from its most frequent OpenAlex topic. Multidisciplinary for journals the Norwegian Register for Scientific Journals (NPI) files under Interdisciplinary Natural Sciences. Empty when OpenAlex has no classified works for the journal. |
| scientific_subfield | Subfield | string | external | Journal-level OpenAlex subfield. For journals classified by their articles (see scientific_field), the most common subfield among those articles within scientific_field; for journals that keep the v3.0.0 classification, the subfield of their most frequent OpenAlex topic. Empty for Multidisciplinary journals. |
| scientific_topic | Topic | string | external | Journal-level OpenAlex topic. For journals classified by their articles (see scientific_field), the most common topic among those articles within scientific_subfield; for journals that keep the v3.0.0 classification, their most frequent OpenAlex topic. Empty for Multidisciplinary journals. |
Journal identifiers (2 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| openalex_source_id | OpenAlex source ID | string | external | OpenAlex Source identifier for the journal, resolved from the journal title and/or ISSN. In v4.0.0, 152 journals whose recorded source OpenAlex had since retired were moved to the current source for their ISSN-L, and their four oa_ journal metrics were refreshed from it in September 2026; the other journals keep the metrics from the original enrichment. |
| issn_l | ISSN-L | string | external | Linking ISSN (the ISSN-L groups print and electronic variants into one identifier) from OpenAlex. |
Journal metrics (7 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| oa_2yr_mean_citedness | Mean citedness (2yr) | float | external | OpenAlex 2-year mean citedness: average citations received by articles published in the last 2 years. Open-data analogue of the Clarivate Journal Impact Factor but computed from OpenAlex citation graph. |
| oa_journal_h_index | Journal h-index | int | external | Journal-level h-index from OpenAlex. |
| oa_journal_works_count | Journal works count | int | external | Total number of works published in the journal (OpenAlex count). |
| oa_journal_cited_by_count | Journal citations | int | external | Total citations received by the journal (OpenAlex count). |
| is_in_doaj | In DOAJ | bool | external | True if the journal is listed in the Directory of Open Access Journals (DOAJ). |
| is_oa | Open access | bool | external | True if OpenAlex classifies the journal as open access. |
| oa_impact_quartile | OpenAlex citedness quartile (per field) | string | computed | Q1 to Q4, computed locally by this project within each scientific_field from oa_2yr_mean_citedness. A journal's position is the share of journals in the same field with strictly lower citedness: Q4 below 0.25, Q3 below 0.50, Q2 below 0.75, Q1 from 0.75. Q1 is therefore roughly the top quarter of the field, so the value accounts for differing citation norms across disciplines, and journals with equal citedness share a quartile. Empty when the journal has no field or no citedness, or its field has fewer than 4 such journals. Releases up to v3.0.0 published one split across all journals despite this description; see CHANGELOG. This is NOT the Clarivate JIF or the Scopus CiteScore quartile. |
Editor bibliometrics (7 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| h_index | h-index | int | external | Author h-index from OpenAlex, looked up by (name, ror_id) pair. |
| total_publications | Publications | int | external | Total number of works by this author in OpenAlex. |
| total_citations | Citations | int | external | Total citations received by this author in OpenAlex. |
| academic_age | Academic age | int | computed | Years since this author's first OpenAlex-indexed publication (current year minus earliest publication year). |
| orcid_source | ORCID source | string | computed | Provenance of the ORCID iD. Values: 'scraped' (the iD was published on the editorial board page itself), 'missing' (no iD resolved for this record), and 'cleared_cross_institution_backfill' (a previously backfilled iD that was withdrawn in v2.x because the candidate profile resolved to a different institution than the editor's, and so could not be trusted). |
| openalex_author_id | OpenAlex author ID | string | external | OpenAlex Author identifier for the matched profile, restoring direct interoperability with the OpenAlex graph. Populated on 531,437 records (57.6%), which is 98.7% of the 538,550 records that matched an OpenAlex author. Added in v3.0.0; earlier releases carried the derived bibliometrics (h_index and friends) without exposing the identifier they came from. |
| oa_match_method | OpenAlex match method | string | computed | How the OpenAlex author profile was resolved, so users can filter by match strength. Values: 'orcid' (matched on ORCID iD, the strongest signal, 344,957 records), 'name_ror' (matched on name plus the ROR-resolved institution, 107,173), 'name_only' (matched on name alone and therefore the weakest tier, most exposed to homonym error, 86,420), and 'unmatched' (no profile found, 383,547). Added in v3.0.0. |
Indexing (7 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| indexed_pubmed | PubMed | bool | external | True if the journal is indexed in PubMed/MEDLINE (matched on ISSN against the NLM catalog). |
| indexed_scopus | Scopus | bool | external | True if the journal is indexed in Scopus (matched on ISSN against the Scopus source list). |
| indexed_wos | Web of Science | bool | external | True if the journal is indexed in Web of Science (matched on ISSN against the WoS master journal list). |
| indexed_doaj | DOAJ | bool | external | True if the journal is listed in the Directory of Open Access Journals (DOAJ). |
| indexed_cope | COPE | bool | external | True if the publisher is a member of COPE (Committee on Publication Ethics). Publisher-level flag applied to all of that publisher's journals. |
| indexed_npi | NPI | bool | external | True if the journal appears in the Norwegian Publishing Indicator register (at either level 1 or level 2). See the separate 'Norwegian Publishing Indicator' group below for the level and discipline fields. |
| indexing_count | Index count | int | computed | Sum of the indexed_* flags (PubMed, Scopus, WoS, DOAJ, COPE, NPI). Range 0–6. A rough journal-quality proxy independent of citation metrics. Used as the indexing weight in the experimental 'weighted power' score on the Network page. |
Norwegian Publishing Indicator (3 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| npi_level | NPI level | string | external | Norwegian Publishing Indicator level. Level 2 marks the top ~20% of journals in the register; level 1 is the remainder of the register; 0 and X denote entries the register carries without a scoring level. Scope is limited to Nordic-relevant disciplines. Stored as a string; 320 rows carry the float-formatted variants '1.0' and '2.0', which downstream code should fold into '1' and '2'. |
| npi_discipline | NPI discipline | string | external | Broad discipline in the NPI register. |
| npi_field | NPI field | string | external | Specific field in the NPI register. |
Funding (6 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| top_funder_1 | Top funder 1 | string | external | Most common funding organization for articles in this journal (from the OpenAlex Works funder metadata). |
| top_funder_1_count | Funder 1 count | int | external | Number of funded articles from the top funder. |
| top_funder_2 | Top funder 2 | string | external | Second most common funder. |
| top_funder_2_count | Funder 2 count | int | external | Count of funded articles. |
| top_funder_3 | Top funder 3 | string | external | Third most common funder. |
| top_funder_3_count | Funder 3 count | int | external | Count of funded articles. |
Board diversity (6 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| board_size | Board size | int | computed | Total number of editors on this journal's board (number of distinct editor rows sharing the same journal). |
| board_pct_female | % female | float | computed | Percentage (0 to 100) of female editors on this journal's board. The denominator is the RESOLVED gender count (male + female), not total board size, so the value is not artificially depressed on boards where many editors have unknown inferred gender. Fixed in v3.0.0: in v2.6.0 and v2.7.0 this column was constant at 0 because a gender-wipe regression emptied the inputs to the board-statistics stage. Any analysis using board_pct_female from those releases must be re-run. |
| board_country_count | Countries on board | int | computed | Number of distinct ror_country values on the board. |
| board_country_hhi | Country HHI | float | computed | Herfindahl–Hirschman Index of country concentration. Sum of squared country shares on the board. 0 = maximally diverse across many countries; 1 = all editors from a single country. |
| board_institution_count | Institutions on board | int | computed | Number of distinct ror_id values on the board. |
| board_mean_h_index | Mean board h-index | float | computed | Arithmetic mean of h_index across board members with a resolved OpenAlex profile. |
Multi-board (3 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| boards_count | Boards served | int | computed | Number of distinct editorial boards this editor serves on in the dataset. |
| publishers_count | Publishers served | int | computed | Number of distinct publishers this editor serves across. |
| is_multi_board | Multi-board | bool | computed | True if boards_count >= 2. |
Metadata (7 columns)
| Column | Display name | Type | Source | Description |
|---|---|---|---|---|
| name_script | Script | string | inferred | Detected writing script of the editor's name (Latin, CJK, Cyrillic, Arabic, etc.). Used to document coverage gaps in downstream inference steps. |
| name_script_region | Script region | string | inferred | Geographic region associated with the name script. Heuristic, not authoritative. |
| name_is_initials | Name is initials | bool | computed | True when the editor's given name is recorded only as initials (for example 'J. R. Smith'), which blocks first-name-based gender inference and weakens name-based OpenAlex matching. True on 26,763 records (2.9%). Added in v3.0.0 so these records can be excluded from name-sensitive analyses. |
| orcid_n_names | Distinct names per ORCID | int | computed | Number of distinct editor name strings that map to this record's ORCID iD in the OpenAlex snapshot. 1 for a clean one-to-one iD. Supporting column for orcid_shared_flag. Added in v3.0.0. |
| orcid_shared_flag | Shared ORCID | bool | computed | True when one ORCID iD maps to several distinct name strings in the OpenAlex snapshot, which usually signals a data-entry error at the source rather than a genuine identity. True on 239 records. Added in v3.0.0 as a transparency flag; these records are retained, not removed. |
| data_version | Version | float | computed | Dataset version stamp for the row, matching CHANGELOG.md. 2026.4 for the v4.0.1 release (v4.0.0: 2026.3). Stored as a number, not a string. |
| enriched_at | Enriched at | datetime | computed | ISO timestamp of when enrichment completed for this row. |