[](#chatglm3-6b)ChatGLM3-6B
===========================

 [Github Repo](https://github.com/THUDM/ChatGLM)   [Twitter](https://twitter.com/thukeg)   [\[GLM@ACL 22\]](https://arxiv.org/abs/2103.10360) [\[GitHub\]](https://github.com/THUDM/GLM)   [\[GLM-130B@ICLR 23\]](https://arxiv.org/abs/2210.02414) [\[GitHub\]](https://github.com/THUDM/GLM-130B)  

 Join our [Slack](https://join.slack.com/t/chatglm/shared_invite/zt-25ti5uohv-A_hs~am_D3Q8XPZMpj7wwQ) and [WeChat](https://github.com/THUDM/ChatGLM/blob/main/resources/WECHAT.md)

Experience the larger-scale ChatGLM model at [chatglm.cn](https://www.chatglm.cn)

[](#-introduction) (Introduction)
-------------------------------------

ChatGLM3-6B  ChatGLM ChatGLM3-6B 

1.  **** ChatGLM3-6B  ChatGLM3-6B-Base ChatGLM3-6B-Base  10B 
2.  **** ChatGLM3-6B  [Prompt ](https://github.com/THUDM/ChatGLM3/blob/main/PROMPT.md)[](https://github.com/THUDM/ChatGLM3/blob/main/tool_using/README.md)Function CallCode Interpreter Agent 
3.  ****  ChatGLM3-6B  ChatGLM-6B-Base ChatGLM3-6B-32K****[](https://open.bigmodel.cn/mla/form)****

ChatGLM3-6B is the latest open-source model in the ChatGLM series. While retaining many excellent features such as smooth dialogue and low deployment threshold from the previous two generations, ChatGLM3-6B introduces the following features:

1.  **More Powerful Base Model:** The base model of ChatGLM3-6B, ChatGLM3-6B-Base, employs a more diverse training dataset, more sufficient training steps, and a more reasonable training strategy. Evaluations on datasets such as semantics, mathematics, reasoning, code, knowledge, etc., show that ChatGLM3-6B-Base has the strongest performance among pre-trained models under 10B.
2.  **More Comprehensive Function Support:** ChatGLM3-6B adopts a newly designed [Prompt format](https://github.com/THUDM/ChatGLM3/blob/main/PROMPT_en.md), in addition to the normal multi-turn dialogue. It also natively supports [function call](https://github.com/THUDM/ChatGLM3/blob/main/tool_using/README_en.md), code interpreter, and complex scenarios such as agent tasks.
3.  **More Comprehensive Open-source Series:** In addition to the dialogue model ChatGLM3-6B, the base model ChatGLM-6B-Base and the long-text dialogue model ChatGLM3-6B-32K are also open-sourced. All the weights are **fully open** for academic research, and after completing the [questionnaire](https://open.bigmodel.cn/mla/form) registration, they are also **allowed for free commercial use**.

[](#-dependencies) (Dependencies)
-----------------------------------------

    pip install protobuf transformers==4.30.2 cpm_kernels torch>=2.0 gradio mdtex2html sentencepiece accelerate
    

[](#-code-usage) (Code Usage)
-------------------------------------

 ChatGLM3-6B 

You can generate dialogue by invoking the ChatGLM3-6B model with the following code:

    >>> from transformers import AutoTokenizer, AutoModel
    >>> tokenizer = AutoTokenizer.from_pretrained("THUDM/chatglm3-6b", trust_remote_code=True)
    >>> model = AutoModel.from_pretrained("THUDM/chatglm3-6b", trust_remote_code=True).half().cuda()
    >>> model = model.eval()
    >>> response, history = model.chat(tokenizer, "", history=[])
    >>> print(response)
    ! ChatGLM-6B,,
    >>> response, history = model.chat(tokenizer, "", history=history)
    >>> print(response)
    ,:
    
    1. :,,
    2. :,,,
    3. :,,,,,
    4. :,,,
    5. :,,,
    6. :,,,,
    
    ,,
    

 DEMO [Github Repo](https://github.com/THUDM/ChatGLM)

For more instructions, including how to run CLI and web demos, and model quantization, please refer to our [Github Repo](https://github.com/THUDM/ChatGLM).

[](#-license) (License)
---------------------------

 [Apache-2.0](/THUDM/chatglm3-6b/blob/main/LICENSE) ChatGLM3-6B  [Model License](/THUDM/chatglm3-6b/blob/main/MODEL_LICENSE)

The code in this repository is open-sourced under the [Apache-2.0 license](/THUDM/chatglm3-6b/blob/main/LICENSE), while the use of the ChatGLM3-6B model weights needs to comply with the [Model License](/THUDM/chatglm3-6b/blob/main/MODEL_LICENSE).

[](#-citation) (Citation)
-----------------------------



If you find our work helpful, please consider citing the following papers.

    @article{zeng2022glm,
      title={Glm-130b: An open bilingual pre-trained model},
      author={Zeng, Aohan and Liu, Xiao and Du, Zhengxiao and Wang, Zihan and Lai, Hanyu and Ding, Ming and Yang, Zhuoyi and Xu, Yifan and Zheng, Wendi and Xia, Xiao and others},
      journal={arXiv preprint arXiv:2210.02414},
      year={2022}
    }
    

    @inproceedings{du2022glm,
      title={GLM: General Language Model Pretraining with Autoregressive Blank Infilling},
      author={Du, Zhengxiao and Qian, Yujie and Liu, Xiao and Ding, Ming and Qiu, Jiezhong and Yang, Zhilin and Tang, Jie},
      booktitle={Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
      pages={320--335},
      year={2022}
    }

## Model Overview

`ChatGLM3-6B` is the latest open-source model in the ChatGLM series from THUDM. It retains many excellent features from previous generations, such as smooth dialogue and low deployment threshold, while introducing several new capabilities. The base model, `ChatGLM3-6B-Base`, employs a more diverse training dataset, more sufficient training steps, and a more reasonable training strategy, making it one of the strongest pre-trained models under 10B parameters.

In addition to the standard multi-turn dialogue, `ChatGLM3-6B` adopts a newly designed [Prompt format](https://github.com/THUDM/ChatGLM3/blob/main/PROMPT_en.md) that natively supports [function call](https://github.com/THUDM/ChatGLM3/blob/main/tool_using/README_en.md), code interpreter, and complex scenarios such as agent tasks. The open-source series also includes the base model `ChatGLM-6B-Base` and the long-text dialogue model `ChatGLM3-6B-32K`.

## Model Inputs and Outputs

### Inputs
- **Text**: The model takes text input, which can be in the form of a multi-turn dialogue or a prompt for the model to respond to.

### Outputs
- **Text**: The model generates human-readable text in response to the input. This can include dialogue responses, code, or task outputs depending on the prompt.

## Capabilities

`ChatGLM3-6B` is a powerful generative language model capable of engaging in smooth, coherent dialogue while also supporting more advanced functionalities like code generation and task completion. Evaluations show the base model, `ChatGLM3-6B-Base`, has strong performance across a variety of datasets including semantics, mathematics, reasoning, code, and knowledge.

## What Can I Use It For?

`ChatGLM3-6B` is well-suited for a wide range of natural language processing tasks, from chatbots and virtual assistants to code generation and task automation. The model's diverse capabilities mean it could be useful in industries like customer service, education, programming, and research.

Some potential use cases include:

- Building conversational AI agents for customer support or personal assistance
- Generating code snippets or even complete programs based on textual descriptions
- Automating repetitive tasks through the model's ability to interpret and execute instructions
- Enhancing language learning and tutoring applications
- Aiding in research and analysis by summarizing information or drawing insights from text

The open licensing of the model also makes it accessible for academic and non-commercial use.

## Things to Try

One interesting aspect of `ChatGLM3-6B` is its ability to handle complex, multi-step prompts and tasks. Try providing the model with a detailed, multi-part instruction or scenario and see how it responds. For example, you could ask it to write a short story with specific plot points and characters, or to solve a complex problem by breaking it down into a series of subtasks.

Another intriguing possibility is to explore the model's code generation and interpretation capabilities. See if you can prompt it to write a working program in a programming language, or to analyze and explain the functionality of a given code snippet.

By pushing the boundaries of what you ask the model to do, you can gain a better understanding of its true capabilities and limitations. The combination of fluent dialogue and more advanced task-completion skills makes `ChatGLM3-6B` a fascinating model to experiment with.