Instructions to use KBLab/megatron-bert-base-swedish-cased-125k with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use KBLab/megatron-bert-base-swedish-cased-125k with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="KBLab/megatron-bert-base-swedish-cased-125k")# Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("KBLab/megatron-bert-base-swedish-cased-125k") model = AutoModelForMaskedLM.from_pretrained("KBLab/megatron-bert-base-swedish-cased-125k", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Megatron-BERT-base Swedish 125k
This BERT model was trained using the Megatron-LM library. The size of the model is a regular BERT-base with 110M parameters. The model was trained on about 70GB of data, consisting mostly of OSCAR and Swedish newspaper text curated by the National Library of Sweden.
Training was done for 125k training steps. Its sister model used the same setup, but was instead trained for 600k steps.
The model has three sister models trained on the same dataset:
Acknowledgements
We gratefully acknowledge the HPC RIVR consortium (https://www.hpc-rivr.si) and EuroHPC JU (https://eurohpc-ju.europa.eu) for funding this research by providing computing resources of the HPC system Vega at the Institute of Information Science (https://www.izum.si).
- Downloads last month
- 64