Does BERT Rediscover a Classical NLP Pipeline?
Jingcheng Niu, Wenjie Lu and Gerald Penn.
COLING 2022 oral
TL;DR
Does BERT store surface knowledge in its bottom layers, syntactic knowledge in its middle layers, and semantic knowledge in its upper layers? Re-examining Jawahar et al. (2019) and Tenney et al.’s (2019) probes, we find that this pipeline-like separation lacks conclusive empirical support. BERT’s structure is linguistically grounded, but in a way more nuanced than layers alone can explain: our novel probe, GridLoc, also takes into account token positions, training rounds, and random seeds, and detects other, stronger regularities suggesting that layer depth may not be the preferred mode of explanation for BERT’s inner workings.

How to Cite
@inproceedings{niu-etal-2022-bert,
title = "Does {BERT} Rediscover a Classical {NLP} Pipeline?",
author = "Niu, Jingcheng and
Lu, Wenjie and
Penn, Gerald",
booktitle = "Proceedings of the 29th International Conference on Computational Linguistics",
month = oct,
year = "2022",
address = "Gyeongju, Republic of Korea",
publisher = "International Committee on Computational Linguistics"
}