- Sophia Ananadiou, Department of Computer Science, University of Manchester, UK
- Dina Demner-Fushman, Tenure Track Investigator, Biomedical Informatics Branch, Lister Hill National Center for Biomedical Communications
- Hercules Dalianis, Professor in Computer and Systems Science, Stockholm University, Sweden
- Hongfang Liu, Professor of Biomedical Informatics, Mayo Clinic
- Josep Maria Haro Abad, Institut de Recerca Sant Joan de Déu
- Bradley Malin, Accenture Professor of Biomedical Informatics, Biostatistics, and Computer Science, Vanderbilt
- Goran Nenadic, Department of Computer Science, University of Manchester, UK
- Aurélie Névéol, LIMSI-CNRS, Université Paris-Sud, France
- Øystein Nytrø, Department of Computer and Information Science, Norges teknisk-naturvitenskapelige universitet (NTNU)
- Carlos Luis Parra Calderón, Head of Technological Innovation at Virgen del Rocío University Hospital, Institute of Biomedicine of Seville, Spain
- Kirk E. Roberts, School of Biomedical Informatics, University of Texas Health Science Center
- Francisco Javier Sanz Valero, Escuela Nacional de Medicina del Trabajo, Instituto de Salud Carlos III, Spain
- Stefan Schulz, Institute for Medical Informatics, Statistics and Documentation, Medical University of Graz, Austria
- Ashish Tendulkar, Machine Learning Specialist at Google
- Michelle Turner, Assistant Research Professor at Barcelona Institute for Global Health, Secretary-Treasurer International Society for Environmental Epidemiology (ISEE)
- Ozlem Uzuner, George Mason University
- Alfonso Valencia Herrera, Barcelona Supercomputing Center (BSC-CNS), Spain
Annotation Guidelines
The MEDDOPROF corpus was manually annotated by linguist experts following annotation guidelines specifically create for this task. These guidelines contain rules for annotating professions, employment statuses and work-related activities (which were not included in this task) in clinical cases in Spanish. Additionally, they also include some considerations regarding the codification of the annotations to the ESCO and SNOMED-CT taxonomies.
Guidelines were created de novo in three phases:
- First, a zero version of the guidelines was developed after annotating a initial batch of ~200 clinical cases and outlining the main problems and difficulties of the data.
- Second, a stable version of guidelines was reached while annotating sample sets of the MEDDOPROF corpus iteratively until quality control was satisfactory.
- Third, guidelines are iteratively refined as manual annotation continues.
The annotation guidelines are available in Zenodo.
Datasets
The MEDDOPROF corpus has been randomly sampled into two subsets: train and test set.
The complete dataset is available in Zenodo.
Sample set
The sample set is composed of 15 clinical cases extracted from the training set. In order to make the sample set somewhat representative of the corpus, we included cases from four different specialties: radiology, oncology, psychiatry and occupational health.
Download the sample set from Zenodo.
Training set
The training set is composed of 1500 clinical cases (~80% of the corpus).
Download the training set from Zenodo.
Codes Reference List
For task 3 (MEDDOPROF-NORM), a reference list with all valid codes is provided. It is a .tsv file with three columns: code, label and alternative label. Codes from two sources are listed: ESCO and SNOMED-CT (these are preceded by the string ‘SCTID:’ in the list). With a few exceptions, professions are mapped to ESCO, while working statuses and activities are mapped to SNOMED-CT.
Download the codes reference list from Zenodo.
Test set
The test set is composed of 344 clinical cases (~20% of the corpus).
Download the test set from Zenodo.
FAQ
If your question is not listed below, email Martin Krallinger to encargo-pln-life@bsc.es or Salvador Lima López to salvador.limalopez@gmail.com.
Q: What is the goal of the shared task?
The goal is, given a collection of clinical reports, to detect occupations mentions and classify them into professions or employment statuses (Track 1), detect who they are referring to (Track 2) and normalize them (Track 3).
Q: Why should I participate?
Demographic information about patients may reveal key aspects for the diagnosis and treatment of their condition. Professions and employment situations are a good example of demographic variables, and are usually only found in unstructured text. The COVID-19 pandemic has highlighted the relevance of occupations at high risk of contagion such as nurses, doctors, hospital cleaners and shopkeepers and of impact on mental health such as health-workers, retired people and the unemployed. Finding these variables facilitates clustering of patients in risk groups and the implementation of the most suitable treatment and preventive measures.
Q: How do I register?
Fill in the following form: https://docs.google.com/forms/d/e/1FAIpQLSclQgJKfqKZgV3M94VQbKcLpqs3OFw66ZuA84Mjz3aYvD3XrA/viewform
Q: How do I submit the results?
See the Submission page for more info.
Q: Can I use additional training data to improve model performance?
Yes, participants may use any additional training data they have available, as long as they describe it in the working notes. We will ask to summarize such resources in your participant paper.
Q: MEDDOPROF has three tracks. Do I need to participate in all of them?
Sub-tracks are independent and participants may participate in one or two of them.
Q: Which controlled vocabularies are used for normalization?
Both the European Skills, Competences, Qualifications and Occupations (ESCO) classification and SNOMED-CT are used for normalization. With some exceptions, professions are generally mapped to ESCO and employment statuses are mapped to SNOMED-CT.
Task Organizers
MEDDOPROF Shared Task 2021 is organized by:
- Martin Krallinger, Barcelona Supercomputing Center, Spain
- Eulàlia Farre-Maduell, Barcelona Supercomputing Center, Spain
- Antonio Miranda Escalada, Barcelona Supercomputing Center, Spain
- Salvador Lima López, Barcelona Supercomputing Center, Spain
- Vicent Brivá-Iglesias, Dublin City University, Ireland
Description of the corpus
General information
The MEDDOPROF corpus is a collection of 1844 clinical cases from over 20 different specialties annotated with professions and employment statuses. The corpus was annotated by a team composed of linguists and clinical experts following specially prepared annotation guidelines, after several cycles of quality control and annotation consistency analysis before annotating the entire dataset. Figure 1 shows a screenshot of a sample manual annotation generated using the brat annotation tool.

The corpus will be distributed in plain text in UTF8 encoding, where each clinical case would be stored as a single file. These clinical case reports were carefully selected to represent records reflecting as much as possible clinical narrative related to electronic clinical reports, including cases from around 20 different medical specialties (such as infectious diseases (including Covid-19 case reports), cardiology, neurology, oncology, psychiatry, urology, internal medicine, emergency and intensive care medicine, radiology, tropical medicine, …). Figure 2 illustrates an example text snippet corresponding to a short sample record.

Additionally, we will also provide the annotation files comprising the character offsets of the tumor morphology entity mentions in TSV (tab-separated values) BRAT format and a TSV file with each individual mapped to the European Skills, Competences, Qualifications and Occupations (ESCO) classification and SNOMED-CT.
The goal of the MEDDOPROF task is to develop automatic occupation detection systems for Spanish medical texts. These systems should rely on the use of the MEDDOPROF corpus, a high-quality Gold Standard clinical corpus of 3000 records based on a manual annotation process done by human clinical coding experts together with an inter-annotator agreement consistency analysis.
The MEDDOPROF task can be approached as a named entity recognition and normalization task. Participants are encouraged to either propose solutions in one of these directions or to combine both approaches. As well, novel approaches are welcomed.
Corpus format
Track 1 – MEDDOPROF-NER: brat annotation format.

Track 2 – MEDDOPROF-CLASS: brat annotation format.

Track 3 – MEDDOPROF-NORM: we provide a single plain text file per clinical case and a tab-separated file with each mention’s code. SNOMED-CT codes have the prefix ‘SCTID:’.

Evaluation Method
MEDDOPROF Shared Task’s sub-tracks will be evaluated in the following way:
Track A – MEDDOPROF-NER
Submissions will be ranked by Precision, Recall and F1-score for each PROFESION [profession] or SITUACION_LABORAL [working status] mention extracted, where the spans overlap entirely (F-score is the primary metric).
A correct prediction must have the same beginning and ending offsets as the Gold Standard annotation, as well as the same label (PROFESION or SITUACION_LABORAL)
Prediction format: brat annotation files (.ANN) with your predictions.
Track B – MEDDOPROF-CLASS
Submissions will be ranked by Precision, Recall and F1-score for each PACIENTE [patient], FAMILIAR [family member], SANITARIO [health professional] or OTROS [others] mention extracted, where the spans overlap entirely (F-score is the primary metric).
A correct prediction must have the same beginning and ending offsets as the Gold Standard annotation, as well as the same label.
Prediction format: brat annotation files (.ANN) with your predictions.
Track C – MEDDOPROF-NORM
For this track, participants will be provided with a list of unique concept identifiers from the European Skills, Competences, Qualifications and Occupations (ESCO) classification and relevant SNOMED-CT terms. Participants will have to detect PROFESION and SITUACION_LABORAL mentions and map each of them to one of the terms in the list. Then, their mappings will be compared to the manually annotated concept ids and evaluated using F1-score.
Precision, Recall and F1-score will be calculated using the following formula:
Precision (P) = true positives/(true positives + false positives)
Recall (R) = true positives/(true positives + false negatives)
F-score (F1) = 2*((P*R)/(P+R))
Evaluation Library
MEDDOPROF’s evaluation will be done using the official evaluation library, which can be downloaded from GitHub. This library is written in Python 3 and intended to be run via command line:
$ python main.py -g ../gs-data/ner/ -p ../toy-data/ner/ -s ner
$ python main.py -g ../gs-data/class/ -p ../toy-data/class -s class
$ python main.py -g ../gs-data/gs-norm.tsv -p ../toy-data/pred-norm.tsv -c ../meddoprof_valid_codes.tsv.tsv -s norm
For all subtasks, the relevant metrics are precision, recall and f1-score. The latter will be used to decide the award winners.
Submission
Awards
MEDDOPROF will award the top three teams of each sub-task. For each of them, the team with the highest F1-score will be awarded a prize of 600 euros, the second one with a prize of 300 euros, and the third one with a prize of 100 euros.
In order to encourage participants to support open knowledge, the team will receive the full amount of the prize as long as they open source the model and a script to use it by other members of the community in a public repository such as Github. This repository must have all the necessary files to make it work by a third party. If they do not, a 50% deduction will be applied to the prize amount.
In case of a tie, the decision will be made following these criteria:
- First, it will be checked which system has been made publicly available in an online repository.
- Second, in case both teams have made public their systems, the team that has submitted a short paper describing the operation of the system will win.
If all tied teams meet these criteria, the prize money will be divided equally among the teams.
A template README file to be used in the repository can be downloaded from: https://github.com/PlanTL-SANIDAD/shared-task-resource-example/