[](#model-card)Model Card
=========================

![Logo](/BAAI/Bunny-Llama-3-8B-V/resolve/main/icon.png)

 [Technical report](https://arxiv.org/abs/2402.11530) |  [Code](https://github.com/BAAI-DCAI/Bunny) |  [3B Demo](https://wisemodel.cn/spaces/baai/Bunny) |  [8B Demo](https://d61b68ac93656b614f.gradio.live/) |  [GGUF](https://huggingface.co/BAAI/Bunny-Llama-3-8B-V-gguf)

This is Bunny-Llama-3-8B-V.

Bunny is a family of lightweight but powerful multimodal models. It offers multiple plug-and-play vision encoders, like EVA-CLIP, SigLIP and language backbones, including Llama-3-8B, Phi-1.5, StableLM-2, Qwen1.5, MiniCPM and Phi-2. To compensate for the decrease in model size, we construct more informative training data by curated selection from a broader data source.

We provide Bunny-Llama-3-8B-V, which is built upon [SigLIP](https://huggingface.co/google/siglip-so400m-patch14-384) and [Llama-3-8B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct). More details about this model can be found in [GitHub](https://github.com/BAAI-DCAI/Bunny).

[![comparison](/BAAI/Bunny-Llama-3-8B-V/resolve/main/comparison.png)](/BAAI/Bunny-Llama-3-8B-V/blob/main/comparison.png)

[](#quickstart)Quickstart
=========================

Here we show a code snippet to show you how to use the model with transformers.

Before running the snippet, you need to install the following dependencies:

    pip install torch transformers accelerate pillow
    

If the CUDA memory is enough, it would be faster to execute this snippet by setting `CUDA_VISIBLE_DEVICES=0`.

Users especially those in Chinese mainland may want to refer to a HuggingFace [mirror site](https://hf-mirror.com).

    import torch
    import transformers
    from transformers import AutoModelForCausalLM, AutoTokenizer
    from PIL import Image
    import warnings
    
    # disable some warnings
    transformers.logging.set_verbosity_error()
    transformers.logging.disable_progress_bar()
    warnings.filterwarnings('ignore')
    
    # set device
    device = 'cuda'  # or cpu
    torch.set_default_device(device)
    
    # create model
    model = AutoModelForCausalLM.from_pretrained(
        'BAAI/Bunny-Llama-3-8B-V',
        torch_dtype=torch.float16, # float32 for cpu
        device_map='auto',
        trust_remote_code=True)
    tokenizer = AutoTokenizer.from_pretrained(
        'BAAI/Bunny-Llama-3-8B-V',
        trust_remote_code=True)
    
    # text prompt
    prompt = 'Why is the image funny?'
    text = f"A chat between a curious user and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions. USER: <image>\n{prompt} ASSISTANT:"
    text_chunks = [tokenizer(chunk).input_ids for chunk in text.split('<image>')]
    input_ids = torch.tensor(text_chunks[0] + [-200] + text_chunks[1][1:], dtype=torch.long).unsqueeze(0).to(device)
    
    # image, sample images can be found in images folder
    image = Image.open('example_2.png')
    image_tensor = model.process_images([image], model.config).to(dtype=model.dtype, device=device)
    
    # generate
    output_ids = model.generate(
        input_ids,
        images=image_tensor,
        max_new_tokens=100,
        use_cache=True)[0]
    
    print(tokenizer.decode(output_ids[input_ids.shape[1]:], skip_special_tokens=True).strip())

## Model Overview

`Bunny-Llama-3-8B-V` is a family of lightweight but powerful multimodal models developed by BAAI. It offers multiple plug-and-play vision encoders, like [EVA-CLIP](https://huggingface.co/google/siglip-so400m-patch14-384) and [SigLIP](https://huggingface.co/google/siglip-so400m-patch14-384), as well as language backbones including [Llama-3-8B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct), [Phi-1.5](https://huggingface.co/decapoda-research/llama-1.5b-hf), [StableLM-2](https://huggingface.co/stabilityai/stablelm-base-alpha-3b), [Qwen1.5](https://huggingface.co/decapoda-research/llama-1.5b-hf), [MiniCPM](https://huggingface.co/decapoda-research/llama-1.5b-hf), and [Phi-2](https://huggingface.co/decapoda-research/llama-1.5b-hf).

## Model Inputs and Outputs

`Bunny-Llama-3-8B-V` is a multimodal model that can consume both text and images, and produce text outputs. 

### Inputs
- **Text Prompt**: A text prompt or instruction that the model uses to generate a response.
- **Image**: An optional image that the model can use to inform its text generation.

### Outputs
- **Generated Text**: The model's response to the provided text prompt and/or image.

## Capabilities

The `Bunny-Llama-3-8B-V` model is capable of generating coherent and relevant text outputs based on a given text prompt and/or image. It can be used for a variety of multimodal tasks, such as image captioning, visual question answering, and image-grounded text generation.

## What Can I Use It For?

`Bunny-Llama-3-8B-V` can be used for a variety of multimodal applications, such as:

- **Image Captioning**: Generate descriptive captions for images.
- **Visual Question Answering**: Answer questions about the contents of an image.
- **Image-Grounded Dialogue**: Generate responses in a conversation that are informed by a relevant image.
- **Multimodal Content Creation**: Produce text outputs that are coherently grounded in visual information.

## Things to Try

Some interesting things to try with `Bunny-Llama-3-8B-V` could include:

- Experimenting with different text prompts and image inputs to see how the model responds.
- Evaluating the model's performance on standard multimodal benchmarks like VQAv2, OKVQA, and COCO Captions.
- Exploring the model's ability to reason about and describe diagrams, charts, and other types of visual information.
- Investigating how the model's performance varies when using different language backbones and vision encoders.