Establishing Strong Baselines for TripClick Health Retrieval; ECIR 2022

Last update: Nov 03, 2022

Overview

TripClick Baselines with Improved Training Data

Welcome 🙌 to the hub-repo of our paper:

Establishing Strong Baselines for TripClick Health Retrieval Sebastian Hofstätter, Sophia Althammer, Mete Sertkan and Allan Hanbury

https://arxiv.org/abs/2201.00365

tl;dr We create strong re-ranking and dense retrieval baselines (BERT_CAT, BERT_DOT, ColBERT, and TK) for TripClick (health ad-hoc retrieval). We improve the – originally too noisy – training data with a simple negative sampling policy. We achieve large gains over BM25 in the re-ranking and retrieval setting on TripClick, which were not achieved with the original baselines. We publish the improved training files for everyone to use.

If you have any questions, suggestions, or want to collaborate please don't hesitate to get in contact with us via Twitter or mail to [email protected]

Please cite our work as:

@misc{hofstaetter2022tripclick,
      title={Establishing Strong Baselines for TripClick Health Retrieval}, 
      author={Sebastian Hofst{\"a}tter and Sophia Althammer and Mete Sertkan and Allan Hanbury},
      year={2022},
      eprint={2201.00365},
      archivePrefix={arXiv},
      primaryClass={cs.IR}
}

Training Files

We publish the improved training files without the text content instead using the ids from TripClick (with permission from the TripClick owners); for the text content please get the full TripClick dataset from the TripClick Github page.

Our training files have the format query_id pos_passage_id neg_passage_id (with tab separation) and are available as a HuggingFace dataset: https://huggingface.co/datasets/sebastian-hofstaetter/tripclick-training

Source Code

The full source-code for our paper is here, as part of our matchmaker library: https://github.com/sebastian-hofstaetter/matchmaker

We provide getting started guides for training re-ranking and retrieval models, as well as a range of evaluation setups.

Pre-Trained Models

Unfortunately, the license of TripClick does not allow us to publish the trained models.

TripClick Baselines Results

For more information and commentary on the results, please see our ECIR paper.

BM25 Top200 Re-Ranking

Model	BERT Instance	HEAD		TORSO		TAIL
		nDCG	MRR	nDCG	MRR	nDCG	MRR
Original Baselines
BM25	--	.140	.276	.206	.283	.267	.258
ConvKNRM	--	.198	.420	.243	.347	.271	.265
TK	--	.208	.434	.272	.381	.295	.280
Our Improved Baselines
TK	--	.232	.472	.300	.390	.345	.319
ColBERT	SciBERT	.270	.556	.326	.426	.374	.347
	PubMedBERT-Abstract	.278	.557	.340	.431	.387	.361
BERT_CAT	DistilBERT	.272	.556	.333	.427	.381	.355
	BERT-Base	.287	.579	.349	.453	.396	.366
	SciBERT	.294	.595	.360	.459	.408	.377
	PubMedBERT-Full	.298	.582	.365	.462	.412	.381
	PubMedBERT-Abstract	.296	.587	.359	.456	.409	.380
Ensemble (Last 3 BERT_CAT)		.303	.601	.370	.472	.420	.392

Dense Retrieval Results

Model	BERT Instance	Head(DCTR)
		[email protected]	[email protected]	[email protected]	[email protected]	[email protected]	[email protected]
Original Baselines
BM25	--	31%	.140	.276	.499	.621	.834
Our Improved Baselines
BERT_DOT	DistilBERT	39%	.236	.512	.550	.648	.813
	SciBERT	41%	.243	.530	.562	.640	.793
	PubMedBERT	40%	.235	.509	.582	.673	.828

Establishing Strong Baselines for TripClick Health Retrieval; ECIR 2022

Related tags

Overview

TripClick Baselines with Improved Training Data

Training Files

Source Code

Pre-Trained Models

TripClick Baselines Results

BM25 Top200 Re-Ranking

Dense Retrieval Results

Owner

Sebastian Hofstätter

Probabilistic Tracklet Scoring and Inpainting for Multiple Object Tracking

Rewrite ultralytics/yolov5 v6.0 opencv inference code based on numpy, no need to rely on pytorch

RodoSol-ALPR Dataset

Study of human inductive biases in CNNs and Transformers.

Exploit Camera Raw Data for Video Super-Resolution via Hidden Markov Model Inference

SFD implement with pytorch

Breast-Cancer-Prediction

Unofficial implementation of "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows" (https://arxiv.org/abs/2103.14030)

[ICML'21] Estimate the accuracy of the classifier in various environments through self-supervision

This program will stylize your photos with fast neural style transfer.

[TOG 2021] PyTorch implementation for the paper: SofGAN: A Portrait Image Generator with Dynamic Styling.

PyTorch implemention of ICCV'21 paper SGPA: Structure-Guided Prior Adaptation for Category-Level 6D Object Pose Estimation

NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions (CVPR2021)

Mmdet benchmark with python

InsCLR: Improving Instance Retrieval with Self-Supervision

Solving Zero-Shot Learning in Named Entity Recognition with Common Sense Knowledge

DeepSpamReview: Detection of Fake Reviews on Online Review Platforms using Deep Learning Architectures. Summer Internship project at CoreView Systems.

RE3: State Entropy Maximization with Random Encoders for Efficient Exploration

A collection of 100 Deep Learning images and visualizations

EsViT: Efficient self-supervised Vision Transformers