![](/LumiOpen/Poro-34B/resolve/main/poro-logo.png)

[](#poro-34b-model-card)Poro 34B Model Card
===========================================

Poro is a 34B parameter decoder-only transformer pretrained on Finnish, English and code. It was trained on 1 trillion tokens. Poro is a fully open source model and is made available under the Apache 2.0 License.

Poro was created in a collaboration between [SiloGen](https://www.silo.ai/silogen) from [Silo AI](https://www.silo.ai/), the [TurkuNLP group](https://turkunlp.org/) of the University of Turku, and [High Performance Language Technologies](https://hplt-project.org/) (HPLT). Training was conducted on the [LUMI supercomputer](https://www.lumi-supercomputer.eu/), using compute resources generously provided by [CSC](https://csc.fi/) - IT Center for Science, Finland.

This project is part of an ongoing effort to create open source large language models for non-English and especially low resource languages like Finnish. Through the combination of English and Finnish training data we get a model that outperforms previous Finnish only models, while also being fluent in English and code, and capable of basic translation between English and Finnish.

Poro 34B is only the first model of our model family. Work is already underway on our next models which will support additional languages, and include features like flash attention, rotary embeddings, and grouped query attention.

_What does Poro mean?_ Poro is the Finnish word for Reindeer!  These animals are native to Finland and hold a significant and historical role in Finnish culture.

[](#model-overview)Model Overview
---------------------------------

_**NOTE:** In addition to being an early research release, Poro is a base model which needs further fine tuning for most use cases._

Poro is a generative pretrained transformer using a BLOOM architecture, and makes use of ALiBi embeddings to support context length extrapolation at inference time.

Hyperparameter

Value

n\_parameters

34.2B

n\_layers

54

n\_heads

56

d\_model

7168

vocab\_size

128000

sequence\_length

2048

[](#poro-research-checkpoints)Poro Research Checkpoints
-------------------------------------------------------

Checkpoints are available as branches in the repository. Checkpoints will be released roughly every 100B tokens. The main branch will always point to the latest checkpoint. The following checkpoints are available:

*   [100B](https://huggingface.co/LumiOpen/Poro-34B/tree/100B)
*   [200B](https://huggingface.co/LumiOpen/Poro-34B/tree/200B)
*   [300B](https://huggingface.co/LumiOpen/Poro-34B/tree/300B)
*   [400B](https://huggingface.co/LumiOpen/Poro-34B/tree/400B)
*   [500B](https://huggingface.co/LumiOpen/Poro-34B/tree/500B)
*   [600B](https://huggingface.co/LumiOpen/Poro-34B/tree/600B)
*   [700B](https://huggingface.co/LumiOpen/Poro-34B/tree/700B)
*   [800B](https://huggingface.co/LumiOpen/Poro-34B/tree/800B)
*   [900B](https://huggingface.co/LumiOpen/Poro-34B/tree/900B)
*   [1000B](https://huggingface.co/LumiOpen/Poro-34B/tree/1000B)

The transformers library allows you to load a checkpoint from a branch as follows:

    branch = "200B"
    model = transformers.AutoModelForCausalLM.from_pretrained(
        "LumiOpen/Poro-34B",
        torch_dtype=torch.bfloat16,
        revision=branch,
    )
    

[](#training)Training
---------------------

Poro was trained on the LUMI supercomputer, using 512 AMD MI250X GPUs. Each MI250X GPU has two Graphics Complex Dies (GCDs) for a world size of 1024 during training, using activation checkpointing, a micro batch size of 1, gradient accumulation of 16, and a 3D parallelism strategy of TP=2, PP=4, DP=128.

Training began in September 2023 using a custom fork of the Megatron-Deepspeed framework. Our code is available [here](https://github.com/TurkuNLP/Megatron-DeepSpeed).

[](#training-hyperparameters)Training Hyperparameters
-----------------------------------------------------

Hyperparameter

Value

Comment

Precision

bfloat16

Optimizer

AdamW

Learning rate

1.5e-4

10B tokens warm-up, cosine decay to 2e-5

Weight decay

1e-1

Batch size

2048

2048 samples x 2048 tokens = 4194304 tokens

[](#tokenizer)Tokenizer
-----------------------

Poro uses a custom 128K Bloom tokenizer trained on the same English, Finnish and Code dataset used to train the model.

[](#dataset)Dataset
-------------------

Poro is being trained on a 1 trillion token mixed dataset of English, Finnish and Code.

Dataset

Notes

Percentage

Epochs

Tokens

SlimPajama

Excluding books3 data

54.16%

1x

541.7B

Finnish

TurkuNLP Finnish dataset

13.05%

4x

131.5B

Tatoeba

English/Finnish sentence pairs

0.81%

1x

8.0B

Starcoder

31.53%

1.52x

315.4B

Project Gutenberg

from Dolma dataset

0.46%

1x

4.5B

The Finnish dataset is a combination of many Finnish resources:

*   [Finnish Internet Parsebank](https://turkunlp.org/finnish_nlp.html)
*   [mC4 multilingual colossal, cleaned Common Crawl](https://huggingface.co/datasets/mc4)
*   [Common Crawl Finnish](https://github.com/turkunlp/CC-Fi)
*   [Finnish Wikipedia](https://fi.wikipedia.org/wiki)
*   [Lnnrot Projekti Lnnrot](http://www.lonnrot.net/)
*   [Suomi24 The Suomi 24 Corpus 2001-2020](http://urn.fi/urn:nbn:fi:lb-2021101527)
*   [Reddit r/Suomi submissions and comments](https://www.reddit.com/r/Suomi)
*   [STT Finnish News Agency Archive 1992-2018](http://urn.fi/urn:nbn:fi:lb-2019041501)
*   [Yle Finnish News Archive 2011-2018](http://urn.fi/urn:nbn:fi:lb-2017070501)
*   [Yle Finnish News Archive 2019-2020](http://urn.fi/urn:nbn:fi:lb-2021050401)
*   [Yle News Archive Easy-to-read Finnish 2011-2018](http://urn.fi/urn:nbn:fi:lb-2019050901)
*   [Yle News Archive Easy-to-read Finnish 2019-2020](http://urn.fi/urn:nbn:fi:lb-2021050701)

[](#evaluation-results)Evaluation Results
-----------------------------------------

Full evaluations for each checkpoint are available on our [Github repo](https://github.com/LumiOpen/evaluation/).

[](#ethical-considerations-and-limitations)Ethical Considerations and Limitations
---------------------------------------------------------------------------------

Poro is an advanced language model, primarily optimized for English, Finnish and code, with no meaningful proficiency in any other languages. As with most AI-driven systems, Poro is a product of the vast data it has been trained on, which may reflect the imperfections, biases, and idiosyncrasies of the wider web. Poro may, at times, produce outputs that can be considered inaccurate, prejudiced, or controversial. Users and developers engaging with Poro should exercise discretion and consider additional evaluation and customization to ensure the model's responses align with their specific needs and ethical standards.

[](#license)License
-------------------

Poro is released under the Apache 2.0 license.

[](#citation)Citation
---------------------

    @misc{luukkonen2024poro,
          title={Poro 34B and the Blessing of Multilinguality}, 
          author={Risto Luukkonen and Jonathan Burdge and Elaine Zosa and Aarne
    Talman and Ville Komulainen and Vin Hatanp and Peter Sarlin and Sampo
    Pyysalo},
          year={2024},
          eprint={2404.01856},
          archivePrefix={arXiv},
          primaryClass={cs.CL}
    }

## Model overview

The `Poro-34B` model is a 34B parameter decoder-only transformer pretrained on Finnish, English, and code by LumiOpen. It was trained on 1 trillion tokens and is fully open source, available under the Apache 2.0 License. Poro outperforms previous Finnish-only models while also being fluent in English and code, and capable of basic translation between English and Finnish. Similar models like [bloom-7b1](https://aimodels.fyi/models/huggingFace/bloom-7b1-bigscience) from BigScience are also large open-source multilingual language models, but Poro is specifically focused on Finnish and English.

## Model inputs and outputs

The `Poro-34B` model is a text-to-text transformer that can be used for a variety of language tasks. It takes raw text as input and generates coherent text as output. The model has a large 128,000 word vocabulary covering Finnish, English, and programming languages.

### Inputs
- Raw text in Finnish, English, or code
- Prompts for specific language tasks like translation, summarization, or generation

### Outputs
- Coherent text in Finnish, English, or code
- Translations between Finnish and English
- Summaries of input text
- Continuations of prompts for generation tasks

## Capabilities

The `Poro-34B` model excels at Finnish and English language understanding and generation. It can fluently write in both languages, perform basic translation, and even generate simple code. The model's large parameter size and training on a huge corpus of data give it strong general language abilities.

## What can I use it for?

The `Poro-34B` model could be used for a variety of natural language processing tasks in Finnish and English. Some potential applications include:

- Content generation: Writing articles, stories, or other text in Finnish or English
- Translation: Translating between Finnish and English
- Language understanding: Answering questions or completing tasks based on Finnish or English text
- Code generation: Generating simple code snippets in various programming languages

The model's open-source nature and strong performance make it a useful tool for researchers and developers working on Finnish and English language AI projects.

## Things to try

One interesting aspect of the `Poro-34B` model is its ability to handle code along with natural language. You could try prompting the model with a programming task in English and see if it can generate the corresponding code. Or prompt it with Finnish text and see if it can accurately translate it to English. The model's large vocabulary and training on diverse data give it fascinating language understanding capabilities to explore.