[](#codeberta)CodeBERTa
=======================

CodeBERTa is a RoBERTa-like model trained on the [CodeSearchNet](https://github.blog/2019-09-26-introducing-the-codesearchnet-challenge/) dataset from GitHub.

Supported languages:

    "go"
    "java"
    "javascript"
    "php"
    "python"
    "ruby"
    

The **tokenizer** is a Byte-level BPE tokenizer trained on the corpus using Hugging Face `tokenizers`.

Because it is trained on a corpus of code (vs. natural language), it encodes the corpus efficiently (the sequences are between 33% to 50% shorter, compared to the same corpus tokenized by gpt2/roberta).

The (small) **model** is a 6-layer, 84M parameters, RoBERTa-like Transformer model  thats the same number of layers & heads as DistilBERT  initialized from the default initialization settings and trained from scratch on the full corpus (~2M functions) for 5 epochs.

### [](#tensorboard-for-this-training-)Tensorboard for this training 

[![tb](https://cdn-media.huggingface.co/CodeBERTa/tensorboard.png)](https://tensorboard.dev/experiment/irRI7jXGQlqmlxXS0I07ew/#scalars)

[](#quick-start-masked-language-modeling-prediction)Quick start: masked language modeling prediction
----------------------------------------------------------------------------------------------------

    PHP_CODE = """
    public static <mask> set(string $key, $value) {
        if (!in_array($key, self::$allowedKeys)) {
            throw new \InvalidArgumentException('Invalid key given');
        }
        self::$storedValues[$key] = $value;
    }
    """.lstrip()
    

### [](#does-the-model-know-how-to-complete-simple-php-code)Does the model know how to complete simple PHP code?

    from transformers import pipeline
    
    fill_mask = pipeline(
        "fill-mask",
        model="huggingface/CodeBERTa-small-v1",
        tokenizer="huggingface/CodeBERTa-small-v1"
    )
    
    fill_mask(PHP_CODE)
    
    ## Top 5 predictions:
    # 
    ' function' # prob 0.9999827146530151
    'function'  # 
    ' void'     # 
    ' def'      # 
    ' final'    # 
    

### [](#yes-that-was-easy--what-about-some-python-warning-this-is-going-to-be-meta)Yes! That was easy  What about some Python (warning: this is going to be meta)

    PYTHON_CODE = """
    def pipeline(
        task: str,
        model: Optional = None,
        framework: Optional[<mask>] = None,
        **kwargs
    ) -> Pipeline:
        pass
    """.lstrip()
    

Results:

    'framework', 'Framework', ' framework', 'None', 'str'
    

> This program can auto-complete itself! 

### [](#just-for-fun-lets-try-to-mask-natural-language-not-code)Just for fun, let's try to mask natural language (not code):

    fill_mask("My name is <mask>.")
    
    # {'sequence': '<s> My name is undefined.</s>', 'score': 0.2548016905784607, 'token': 3353}
    # {'sequence': '<s> My name is required.</s>', 'score': 0.07290805131196976, 'token': 2371}
    # {'sequence': '<s> My name is null.</s>', 'score': 0.06323737651109695, 'token': 469}
    # {'sequence': '<s> My name is name.</s>', 'score': 0.021919190883636475, 'token': 652}
    # {'sequence': '<s> My name is disabled.</s>', 'score': 0.019681859761476517, 'token': 7434}
    

This (kind of) works because code contains comments (which contain natural language).

Of course, the most frequent name for a Computer scientist must be undefined .

[](#downstream-task-programming-language-identification)Downstream task: [programming language identification](https://huggingface.co/huggingface/CodeBERTa-language-id)
------------------------------------------------------------------------------------------------------------------------------------------------------------------------

See the model card for **[`huggingface/CodeBERTa-language-id`](https://huggingface.co/huggingface/CodeBERTa-language-id)** .

  

[](#codesearchnet-citation)CodeSearchNet citation
-------------------------------------------------

@article{husain_codesearchnet_2019,
        title = {{CodeSearchNet} {Challenge}: {Evaluating} the {State} of {Semantic} {Code} {Search}},
        shorttitle = {{CodeSearchNet} {Challenge}},
        url = {http://arxiv.org/abs/1909.09436},
        urldate = {2020-03-12},
        journal = {arXiv:1909.09436 [cs, stat]},
        author = {Husain, Hamel and Wu, Ho-Hsiang and Gazit, Tiferet and Allamanis, Miltiadis and Brockschmidt, Marc},
        month = sep,
        year = {2019},
        note = {arXiv: 1909.09436},
    }

## Model overview

`CodeBERTa-small-v1` is a RoBERTa-like model trained on the [CodeSearchNet](https://github.com/github/CodeSearchNet) dataset from GitHub. The model supports 6 programming languages: Go, Java, JavaScript, PHP, Python, and Ruby. It uses a Byte-level BPE tokenizer trained on the corpus, which results in 33%-50% shorter sequences compared to models like GPT-2 or RoBERTa trained on natural language. The small model has 6 layers, 84M parameters, and is initialized and trained from scratch on the full 2M function corpus for 5 epochs.

Similar AI models include [BERT multilingual base model (uncased)](https://aimodels.fyi/models/huggingFace/bert-base-multilingual-uncased-google-bert), [BERT base model (uncased)](https://aimodels.fyi/models/huggingFace/bert-base-uncased-google-bert), [BERT large model (uncased)](https://aimodels.fyi/models/huggingFace/bert-large-uncased-google-bert), and [RoBERTa large model](https://aimodels.fyi/models/huggingFace/roberta-large-facebookai). These models are also trained on large corpora of text, but target natural language rather than programming code.

## Model inputs and outputs

### Inputs
- Text containing programming code in one of the 6 supported languages (Go, Java, JavaScript, PHP, Python, Ruby)

### Outputs
- Predicted missing token(s) in the input text, given a masked language modeling task
- Representations of the input code that can be used for downstream tasks like code search or classification

## Capabilities

`CodeBERTa-small-v1` is able to understand and complete simple programming code snippets in the 6 supported languages. For example, it can predict the missing `function` keyword in a PHP code snippet. The model also exhibits some metalearning capabilities, being able to complete a Python code snippet that defines the `pipeline` function from the Transformers library.

## What can I use it for?

The `CodeBERTa-small-v1` model can be used for a variety of tasks related to programming code understanding and generation, such as:

- **Code completion**: Suggesting the most likely tokens to complete a partially written code snippet.
- **Code search**: Finding relevant code examples by encoding code into semantic representations.
- **Code classification**: Assigning tags or labels to code based on its content and functionality.

The model is particularly well-suited for applications that involve processing and understanding large codebases, such as code recommendation systems, automated code review tools, or code-related question answering.

## Things to try

One interesting aspect of `CodeBERTa-small-v1` is its ability to handle both code and natural language. You can experiment with using the model to fill in missing words in a mix of code and text, or to generate text that seamlessly incorporates code snippets. This could be useful for tasks like writing programming-related documentation or tutorials.

Another thing to try is fine-tuning the model on a specific programming language or task, to see if you can improve its performance compared to the out-of-the-box capabilities. The small size of the model also makes it a good candidate for deploying on resource-constrained environments like edge devices.