Codebook

Column definitions for the public release dataset. 77 columns across 14 groups.

Provenance

Every column is tagged with one of four provenance flags so you can tell at a glance how the value was produced. Hover the badge on any row for the full derivation method.

collected · 8 Direct from publisher HTML
external · 39 From authoritative API (ROR, OpenAlex, DOAJ, COPE, NPI)
inferred · 7 Algorithmic / ML inference — may carry bias
computed · 23 Derived arithmetically from other columns

Inferred columns (especially gender) carry systematic bias against non-Latin-script names — see the methodology page for coverage details.

Core (11 columns)

Column Display name Type Source Description
publisher Publisher string collected Publisher name (e.g., Elsevier, Springer Nature, Wiley), known from the collection configuration for each source.
journal Journal string collected Journal title as listed on the publisher website.
editor Name string collected Editor full name as it appeared on the editorial board page, after mojibake repair and light cleaning.
first_name First name string computed Given name extracted from `editor` via the probablepeople ML name parser. Used as the lookup key for WGND 2.0 gender inference; preserved in the public release so users can re-use the parsed value without rerunning probablepeople.
last_name Last name string computed Surname extracted from `editor` via the probablepeople ML name parser. Empty when probablepeople could not parse the string; the original `editor` field is always preserved.
role Role (raw) string collected Role label as listed on the publisher page, unmodified.
role_std Role (standardized) string computed Standardized role, one of exactly ten canonical values: editor_in_chief, senior_editor, deputy_editor, associate_editor, section_editor, reviewing_editor, guest_editor, editorial_board_member, editor, other. The mapping is deterministic from the raw role string. 'senior_editor' was introduced in v3.0.0; 'other' absorbs roughly 6% of positions that do not match the canonical set.
affiliation Affiliation string collected Institutional affiliation as listed, after universal cleaning (strips roles, credentials, dates, junk). Not canonicalised; see ror_name for the authoritative form.
orcid ORCID string collected 16-digit ORCID iD. Preferred from the collected value when available; otherwise backfilled from OpenAlex or the ORCID API (see orcid_source).
source_url Source URL string collected Public publisher URL where this record was collected.
scraped_at Collected at datetime collected ISO-8601 timestamp of when this record was collected. The column name is retained for backward compatibility with earlier releases.

Identity (1 columns)

Column Display name Type Source Description
editor_key Editor key string computed Composite identity key used to deduplicate editors across boards; the basis for the 744,842 unique-editor count. Resolved in priority order and prefixed with the tier that produced it: 'orcid:<orcid>|<ror>' when an ORCID iD is available (346,713 records, 250,217 unique editors); 'ror:<ror_id>|<name>' when there is no iD but the affiliation resolved to ROR (293,742 records, 235,869 unique); 'aff:<affiliation>|<name>' otherwise (281,642 records, 258,756 unique, of which 51,695 have an empty affiliation and so fall back to name alone). Published in v3.0.0 at reviewer request so the unique-editor count is reproducible without reconstructing the key. Because the lowest tier is name-based, distinct people who share a name and affiliation string can still collide there; the ORCID and ROR tiers are the reliable ones.

Gender (5 columns)

Column Display name Type Source Description
gender Gender string inferred Country-aware inference from the editor's first name and ror_country using WGND 2.0 (Raffo & Lax-Martinez, WIPO 2021; Harvard Dataverse DOI 10.7910/DVN/MSEGSJ, roughly 3.5 million names across 195 countries). The same name can resolve differently by country (for example 'Andrea' is typically male in Italy and female in the United States). gender-guesser is a tertiary fallback for names absent from WGND. Values: male, female, unknown. The legacy 'andy' (androgynous) value was retired in v3.0.0. Self-reported gender is NOT available in this dataset; every value is inferred and should be treated as such.
gender_raw Gender (raw) string inferred Backwards-compatibility column kept since v1; mirrors gender exactly (WGND has no mostly_male / mostly_female granularity). Retained so v1-era scripts keep running; new work should read gender.
gender_prob Gender weight float inferred WGND weight in [0, 1]: the probability that an individual with this first name in this country has the inferred gender. Replaces the legacy 3-bucket {0.0, 0.75, 1.0} confidence; downstream filters should use thresholds like `>= 0.95`. For gender-guesser fallback rows, 1.0 = certain, 0.75 = mostly_*, 0.0 = unknown.
gender_nobs WGND sample size int inferred Number of observed individuals in the WGND cell (name × country) underlying the inference. Useful for filtering low-confidence cells in analysis. 0 for gender-guesser-fallback rows.
gender_source Gender source string inferred Provenance of the inference: `wgnd_country` (matched on first_name + ror_country, strongest signal), `wgnd_global` (matched on first_name only, used when country was missing or the country-specific cell was empty), `gender_guesser` (tertiary fallback, country-blind), or `unknown`.

Institution (ROR) (8 columns)

Column Display name Type Source Description
ror_id ROR ID string external Research Organization Registry identifier resolved from the raw affiliation string via the ROR /organizations?affiliation= fuzzy-match endpoint, subject to score and token-overlap guards (see CHANGELOG v2.0.0) that reject low-confidence matches. Seven organisations drew clusters of false matches (issue #30). v4.0.0 re-applied the v2.0.0 repair, which later re-enrichments had undone: a match to one of the seven is kept only when the affiliation names that organisation.
ror_name Institution string external Canonical institution name from the ROR v2 record (ror_display name).
ror_country Country string external Country name from ROR/GeoNames for the primary location of the institution.
ror_city City string external City of the primary location in the ROR record of the row's ror_id (GeoNames). Taken from ROR for every row with a ror_id since v4.0.1, which filled 5,537 rows left empty in v4.0.0.
ror_state State/Province string external State or province of the primary location in the ROR record of the row's ror_id, as ROR gives it: usually a name, occasionally a GeoNames code (e.g. SG.01); empty when the ROR record gives none. Taken from ROR for every row with a ror_id since v4.0.1.
org_type Org type string external Organization type from ROR: education, healthcare, government, facility, nonprofit, company, archive, or other. Taken from the first entry in the ROR record's types array.
latitude Latitude float external Geographic latitude of the institution from ROR/GeoNames.
longitude Longitude float external Geographic longitude of the institution from ROR/GeoNames.

Classification (OpenAlex) (4 columns)

Column Display name Type Source Description
scientific_domain Domain string external OpenAlex domain of scientific_field (Health Sciences, Life Sciences, Physical Sciences or Social Sciences), or Multidisciplinary. Journal-level: identical for every editor of a journal.
scientific_field Field string external Journal-level OpenAlex field (e.g. Medicine, Engineering, Psychology), identical for every editor of a journal: it describes the journal, not the editor's own research. Since v4.0.0 it is the field of most of the journal's own research articles with an abstract, by primary topic (falling back to all articles, then all works, below 30 classified works), provided that field covers at least 40% of those articles or the journal's earlier topic was one of OpenAlex's catch-all topics. Otherwise the journal keeps its v3.0.0 classification, which came from its most frequent OpenAlex topic. Multidisciplinary for journals the Norwegian Register for Scientific Journals (NPI) files under Interdisciplinary Natural Sciences. Empty when OpenAlex has no classified works for the journal.
scientific_subfield Subfield string external Journal-level OpenAlex subfield. For journals classified by their articles (see scientific_field), the most common subfield among those articles within scientific_field; for journals that keep the v3.0.0 classification, the subfield of their most frequent OpenAlex topic. Empty for Multidisciplinary journals.
scientific_topic Topic string external Journal-level OpenAlex topic. For journals classified by their articles (see scientific_field), the most common topic among those articles within scientific_subfield; for journals that keep the v3.0.0 classification, their most frequent OpenAlex topic. Empty for Multidisciplinary journals.

Journal identifiers (2 columns)

Column Display name Type Source Description
openalex_source_id OpenAlex source ID string external OpenAlex Source identifier for the journal, resolved from the journal title and/or ISSN. In v4.0.0, 152 journals whose recorded source OpenAlex had since retired were moved to the current source for their ISSN-L, and their four oa_ journal metrics were refreshed from it in September 2026; the other journals keep the metrics from the original enrichment.
issn_l ISSN-L string external Linking ISSN (the ISSN-L groups print and electronic variants into one identifier) from OpenAlex.

Journal metrics (7 columns)

Column Display name Type Source Description
oa_2yr_mean_citedness Mean citedness (2yr) float external OpenAlex 2-year mean citedness: average citations received by articles published in the last 2 years. Open-data analogue of the Clarivate Journal Impact Factor but computed from OpenAlex citation graph.
oa_journal_h_index Journal h-index int external Journal-level h-index from OpenAlex.
oa_journal_works_count Journal works count int external Total number of works published in the journal (OpenAlex count).
oa_journal_cited_by_count Journal citations int external Total citations received by the journal (OpenAlex count).
is_in_doaj In DOAJ bool external True if the journal is listed in the Directory of Open Access Journals (DOAJ).
is_oa Open access bool external True if OpenAlex classifies the journal as open access.
oa_impact_quartile OpenAlex citedness quartile (per field) string computed Q1 to Q4, computed locally by this project within each scientific_field from oa_2yr_mean_citedness. A journal's position is the share of journals in the same field with strictly lower citedness: Q4 below 0.25, Q3 below 0.50, Q2 below 0.75, Q1 from 0.75. Q1 is therefore roughly the top quarter of the field, so the value accounts for differing citation norms across disciplines, and journals with equal citedness share a quartile. Empty when the journal has no field or no citedness, or its field has fewer than 4 such journals. Releases up to v3.0.0 published one split across all journals despite this description; see CHANGELOG. This is NOT the Clarivate JIF or the Scopus CiteScore quartile.

Editor bibliometrics (7 columns)

Column Display name Type Source Description
h_index h-index int external Author h-index from OpenAlex, looked up by (name, ror_id) pair.
total_publications Publications int external Total number of works by this author in OpenAlex.
total_citations Citations int external Total citations received by this author in OpenAlex.
academic_age Academic age int computed Years since this author's first OpenAlex-indexed publication (current year minus earliest publication year).
orcid_source ORCID source string computed Provenance of the ORCID iD. Values: 'scraped' (the iD was published on the editorial board page itself), 'missing' (no iD resolved for this record), and 'cleared_cross_institution_backfill' (a previously backfilled iD that was withdrawn in v2.x because the candidate profile resolved to a different institution than the editor's, and so could not be trusted).
openalex_author_id OpenAlex author ID string external OpenAlex Author identifier for the matched profile, restoring direct interoperability with the OpenAlex graph. Populated on 531,437 records (57.6%), which is 98.7% of the 538,550 records that matched an OpenAlex author. Added in v3.0.0; earlier releases carried the derived bibliometrics (h_index and friends) without exposing the identifier they came from.
oa_match_method OpenAlex match method string computed How the OpenAlex author profile was resolved, so users can filter by match strength. Values: 'orcid' (matched on ORCID iD, the strongest signal, 344,957 records), 'name_ror' (matched on name plus the ROR-resolved institution, 107,173), 'name_only' (matched on name alone and therefore the weakest tier, most exposed to homonym error, 86,420), and 'unmatched' (no profile found, 383,547). Added in v3.0.0.

Indexing (7 columns)

Column Display name Type Source Description
indexed_pubmed PubMed bool external True if the journal is indexed in PubMed/MEDLINE (matched on ISSN against the NLM catalog).
indexed_scopus Scopus bool external True if the journal is indexed in Scopus (matched on ISSN against the Scopus source list).
indexed_wos Web of Science bool external True if the journal is indexed in Web of Science (matched on ISSN against the WoS master journal list).
indexed_doaj DOAJ bool external True if the journal is listed in the Directory of Open Access Journals (DOAJ).
indexed_cope COPE bool external True if the publisher is a member of COPE (Committee on Publication Ethics). Publisher-level flag applied to all of that publisher's journals.
indexed_npi NPI bool external True if the journal appears in the Norwegian Publishing Indicator register (at either level 1 or level 2). See the separate 'Norwegian Publishing Indicator' group below for the level and discipline fields.
indexing_count Index count int computed Sum of the indexed_* flags (PubMed, Scopus, WoS, DOAJ, COPE, NPI). Range 0–6. A rough journal-quality proxy independent of citation metrics. Used as the indexing weight in the experimental 'weighted power' score on the Network page.

Norwegian Publishing Indicator (3 columns)

Column Display name Type Source Description
npi_level NPI level string external Norwegian Publishing Indicator level. Level 2 marks the top ~20% of journals in the register; level 1 is the remainder of the register; 0 and X denote entries the register carries without a scoring level. Scope is limited to Nordic-relevant disciplines. Stored as a string; 320 rows carry the float-formatted variants '1.0' and '2.0', which downstream code should fold into '1' and '2'.
npi_discipline NPI discipline string external Broad discipline in the NPI register.
npi_field NPI field string external Specific field in the NPI register.

Funding (6 columns)

Column Display name Type Source Description
top_funder_1 Top funder 1 string external Most common funding organization for articles in this journal (from the OpenAlex Works funder metadata).
top_funder_1_count Funder 1 count int external Number of funded articles from the top funder.
top_funder_2 Top funder 2 string external Second most common funder.
top_funder_2_count Funder 2 count int external Count of funded articles.
top_funder_3 Top funder 3 string external Third most common funder.
top_funder_3_count Funder 3 count int external Count of funded articles.

Board diversity (6 columns)

Column Display name Type Source Description
board_size Board size int computed Total number of editors on this journal's board (number of distinct editor rows sharing the same journal).
board_pct_female % female float computed Percentage (0 to 100) of female editors on this journal's board. The denominator is the RESOLVED gender count (male + female), not total board size, so the value is not artificially depressed on boards where many editors have unknown inferred gender. Fixed in v3.0.0: in v2.6.0 and v2.7.0 this column was constant at 0 because a gender-wipe regression emptied the inputs to the board-statistics stage. Any analysis using board_pct_female from those releases must be re-run.
board_country_count Countries on board int computed Number of distinct ror_country values on the board.
board_country_hhi Country HHI float computed Herfindahl–Hirschman Index of country concentration. Sum of squared country shares on the board. 0 = maximally diverse across many countries; 1 = all editors from a single country.
board_institution_count Institutions on board int computed Number of distinct ror_id values on the board.
board_mean_h_index Mean board h-index float computed Arithmetic mean of h_index across board members with a resolved OpenAlex profile.

Multi-board (3 columns)

Column Display name Type Source Description
boards_count Boards served int computed Number of distinct editorial boards this editor serves on in the dataset.
publishers_count Publishers served int computed Number of distinct publishers this editor serves across.
is_multi_board Multi-board bool computed True if boards_count >= 2.

Metadata (7 columns)

Column Display name Type Source Description
name_script Script string inferred Detected writing script of the editor's name (Latin, CJK, Cyrillic, Arabic, etc.). Used to document coverage gaps in downstream inference steps.
name_script_region Script region string inferred Geographic region associated with the name script. Heuristic, not authoritative.
name_is_initials Name is initials bool computed True when the editor's given name is recorded only as initials (for example 'J. R. Smith'), which blocks first-name-based gender inference and weakens name-based OpenAlex matching. True on 26,763 records (2.9%). Added in v3.0.0 so these records can be excluded from name-sensitive analyses.
orcid_n_names Distinct names per ORCID int computed Number of distinct editor name strings that map to this record's ORCID iD in the OpenAlex snapshot. 1 for a clean one-to-one iD. Supporting column for orcid_shared_flag. Added in v3.0.0.
orcid_shared_flag Shared ORCID bool computed True when one ORCID iD maps to several distinct name strings in the OpenAlex snapshot, which usually signals a data-entry error at the source rather than a genuine identity. True on 239 records. Added in v3.0.0 as a transparency flag; these records are retained, not removed.
data_version Version float computed Dataset version stamp for the row, matching CHANGELOG.md. 2026.4 for the v4.0.1 release (v4.0.0: 2026.3). Stored as a number, not a string.
enriched_at Enriched at datetime computed ISO timestamp of when enrichment completed for this row.