jamesthesnake · jamesthesnake · Mar 23, 2023 · Feb 6, 2023 · Feb 6, 2023 · Feb 6, 2023
diff --git a/.compatibility b/.compatibility
@@ -0,0 +1,3 @@
+1.12.0-11.3.0
+1.11.0-11.3.0
+1.10.1-11.3.0
diff --git a/.cuda_ext.json b/.cuda_ext.json
@@ -0,0 +1,16 @@
+{
+  "build": [
+    {
+      "torch_command": "pip install torch==1.12.1+cu102 torchvision==0.13.1+cu102 torchaudio==0.12.1 --extra-index-url https://download.pytorch.org/whl/cu102",
+      "cuda_image": "hpcaitech/cuda-conda:10.2"
+    },
+    {
+      "torch_command": "pip install torch==1.12.1+cu113 torchvision==0.13.1+cu113 torchaudio==0.12.1 --extra-index-url https://download.pytorch.org/whl/cu113",
+      "cuda_image": "hpcaitech/cuda-conda:11.3"
+    },
+    {
+      "torch_command": "pip install torch==1.12.1+cu116 torchvision==0.13.1+cu116 torchaudio==0.12.1 --extra-index-url https://download.pytorch.org/whl/cu116",
+      "cuda_image": "hpcaitech/cuda-conda:11.6"
+    }
+  ]
+}
diff --git a/.github/pull_request_template.md b/.github/pull_request_template.md
@@ -0,0 +1,36 @@
+## 📌 Checklist before creating the PR
+
+- [ ] I have created an issue for this PR for traceability
+- [ ] The title follows the standard format: `[doc/gemini/tensor/...]: A concise description`
+- [ ] I have added relevant tags if possible for us to better distinguish different PRs
+
+
+## 🚨 Issue number
+
+> Link this PR to your issue with words like fixed to automatically close the linked issue upon merge
+>
+> e.g. `fixed #1234`, `closed #1234`, `resolved #1234`
+
+
+
+## 📝 What does this PR do?
+
+> Summarize your work here.
+> if you have any plots/diagrams/screenshots/tables, please attach them here.
+
+
+
+## 💥 Checklist before requesting a review
+
+- [ ] I have linked my PR to an issue ([instruction](https://docs.github.com/en/issues/tracking-your-work-with-issues/linking-a-pull-request-to-an-issue))
+- [ ] My issue clearly describes the problem/feature/proposal, with diagrams/charts/table/code if possible
+- [ ] I have performed a self-review of my code
+- [ ] I have added thorough tests.
+- [ ] I have added docstrings for all the functions/methods I implemented
+
+## ⭐️ Do you enjoy contributing to Colossal-AI?
+
+- [ ] 🌝 Yes, I do.
+- [ ] 🌚 No, I don't.
+
+Tell us more if you don't enjoy contributing to Colossal-AI.
diff --git a/.github/workflows/README.md b/.github/workflows/README.md
@@ -0,0 +1,157 @@
+# CI/CD
+
+## Table of Contents
+
+- [CI/CD](#cicd)
+  - [Table of Contents](#table-of-contents)
+  - [Overview](#overview)
+  - [Workflows](#workflows)
+    - [Code Style Check](#code-style-check)
+    - [Unit Test](#unit-test)
+    - [Example Test](#example-test)
+      - [Example Test on Dispatch](#example-test-on-dispatch)
+    - [Compatibility Test](#compatibility-test)
+      - [Compatibility Test on Dispatch](#compatibility-test-on-dispatch)
+    - [Release](#release)
+    - [User Friendliness](#user-friendliness)
+    - [Commmunity](#commmunity)
+  - [Configuration](#configuration)
+  - [Progress Log](#progress-log)
+
+## Overview
+
+Automation makes our development more efficient as the machine automatically run the pre-defined tasks for the contributors.
+This saves a lot of manual work and allow the developer to fully focus on the features and bug fixes.
+In Colossal-AI, we use [GitHub Actions](https://github.com/features/actions) to automate a wide range of workflows to ensure the robustness of the software.
+In the section below, we will dive into the details of different workflows available.
+
+## Workflows
+
+Refer to this [documentation](https://docs.github.com/en/actions/managing-workflow-runs/manually-running-a-workflow) on how to manually trigger a workflow.
+I will provide the details of each workflow below.
+
+**A PR which changes the `version.txt` is considered as a release PR in the following coontext.**
+
+
+### Code Style Check
+
+| Workflow Name | File name         | Description                                                                                                    |
+| ------------- | ----------------- | -------------------------------------------------------------------------------------------------------------- |
+| `post-commit` | `post_commit.yml` | This workflow runs pre-commit checks for changed files to achieve code style consistency after a PR is merged. |
+
+### Unit Test
+
+| Workflow Name          | File name                  | Description                                                                                                                                       |
+| ---------------------- | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
+| `Build on PR`          | `build_on_pr.yml`          | This workflow is triggered when the label `Run build and Test` is assigned to a PR. It will run all the unit tests in the repository with 4 GPUs. |
+| `Build on Schedule`    | `build_on_schedule.yml`    | This workflow will run the unit tests everyday with 8 GPUs. The result is sent to Lark.                                                           |
+| `Report test coverage` | `report_test_coverage.yml` | This PR will put up a comment to report the test coverage results when `Build` is done.                                                           |
+
+### Example Test
+
+| Workflow Name              | File name                       | Description                                                                    |
+| -------------------------- | ------------------------------- | ------------------------------------------------------------------------------ |
+| `Test example on PR`       | `example_check_on_pr.yml`       | The example will be automatically tested if its files are changed in the PR    |
+| `Test example on Schedule` | `example_check_on_schedule.yml` | This workflow will test all examples every Sunday. The result is sent to Lark. |
+| `Example Test on Dispatch` | `example_check_on_dispatch.yml` | Manually test a specified example.                                             |
+
+#### Example Test on Dispatch
+
+This workflow is triggered by manually dispatching the workflow. It has the following input parameters:
+- `example_directory`: the example directory to test. Multiple directories are supported and must be separated b$$y comma. For example, language/gpt, images/vit. Simply input language or simply gpt does not work.
+
+### Compatibility Test
+
+| Workflow Name                    | File name                            | Description                                                                                                          |
+| -------------------------------- | ------------------------------------ | -------------------------------------------------------------------------------------------------------------------- |
+| `Compatibility Test on PR`       | `compatibility_test_on_pr.yml`       | Check Colossal-AI's compatiblity when `version.txt` is changed in a PR.                                              |
+| `Compatibility Test on Schedule` | `compatibility_test_on_schedule.yml` | This workflow will check the compatiblity of Colossal-AI against PyTorch specified in `.compatibility` every Sunday. |
+| `Compatiblity Test on Dispatch`  | `compatibility_test_on_dispatch.yml` | Test PyTorch Compatibility manually.                                                                                 |
+
+
+#### Compatibility Test on Dispatch
+This workflow is triggered by manually dispatching the workflow. It has the following input parameters:
+- `torch version`:torch version to test against, multiple versions are supported but must be separated by comma. The default is value is all, which will test all available torch versions listed in this [repository](https://github.com/hpcaitech/public_assets/tree/main/colossalai/torch_build/torch_wheels).
+- `cuda version`: cuda versions to test against, multiple versions are supported but must be separated by comma. The CUDA versions must be present in our [DockerHub repository](https://hub.docker.com/r/hpcaitech/cuda-conda).
+
+> It only test the compatiblity of the main branch
+
+
+### Release
+
+| Workflow Name                                   | File name                                   | Description                                                                                                   |
+| ----------------------------------------------- | ------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
+| `Draft GitHub Release Post`                     | `draft_github_release_post_after_merge.yml` | Compose a GitHub release post draft based on the commit history when a release PR is merged.                  |
+| `Publish to PyPI`                               | `release_pypi_after_merge.yml`              | Build and release the wheel to PyPI when a release PR is merged. The result is sent to Lark.                  |
+| `Publish Nightly Version to PyPI`               | `release_nightly_on_schedule.yml`           | Build and release the nightly wheel to PyPI as `colossalai-nightly` every Sunday. The result is sent to Lark. |
+| `Publish Docker Image to DockerHub after Merge` | `release_docker_after_merge.yml`            | Build and release the Docker image to DockerHub when a release PR is merged.  The result is sent to Lark.     |
+| `Check CUDA Extension Build Before Merge`       | `cuda_ext_check_before_merge.yml`           | Build CUDA extensions with different CUDA versions when a release PR is created.                              |
+| `Publish to Test-PyPI Before Merge`             | `release_test_pypi_before_merge.yml`        | Release to test-pypi to simulate user installation when a release PR is created.                              |
+
+
+### User Friendliness
+
+| Workflow Name           | File name               | Description                                                                                                                            |
+| ----------------------- | ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
+| `issue-translate`       | `translate_comment.yml` | This workflow is triggered when a new issue comment is created. The comment will be translated into English if not written in English. |
+| `Synchronize submodule` | `submodule.yml`         | This workflow will check if any git submodule is updated. If so, it will create a PR to update the submodule pointers.                 |
+| `Close inactive issues` | `close_inactive.yml`    | This workflow will close issues which are stale for 14 days.                                                                           |
+
+### Commmunity
+
+| Workflow Name                                | File name                        | Description                                                                      |
+| -------------------------------------------- | -------------------------------- | -------------------------------------------------------------------------------- |
+| `Generate Community Report and Send to Lark` | `report_leaderboard_to_lark.yml` | Collect contribution and user engagement stats and share with Lark every Friday. |
+
+## Configuration
+
+This section lists the files used to configure the workflow.
+
+1. `.compatibility`
+
+This `.compatibility` file is to tell GitHub Actions which PyTorch and CUDA versions to test against. Each line in the file is in the format `${torch-version}-${cuda-version}`, which is a tag for Docker image. Thus, this tag must be present in the [docker registry](https://hub.docker.com/r/pytorch/conda-cuda) so as to perform the test.
+
+2. `.cuda_ext.json`
+
+This file controls which CUDA versions will be checked against CUDA extenson built. You can add a new entry according to the json schema below to check the AOT build of PyTorch extensions before release.
+
+```json
+{
+  "build": [
+    {
+      "torch_command": "",
+      "cuda_image": ""
+    },
+  ]
+}
+```
+
+## Progress Log
+
+- [x] Code style check
+  - [x] post-commit check
+- [x] unit testing
+  - [x] test on PR
+  - [x] report test coverage
+  - [x] regular test
+- [x] release
+  - [x] pypi release
+  - [x] test-pypi simulation
+  - [x] nightly build
+  - [x] docker build
+  - [x] draft release post
+- [x] example check
+  - [x] check on PR
+  - [x] regular check
+  - [x] manual dispatch
+- [x] compatiblity check
+  - [x] check on PR
+  - [x] manual dispatch
+  - [x] auto test when release
+- [x] community
+  - [x] contribution report
+  - [x] user engagement report
+- [x] helpers
+  - [x] comment translation
+  - [x] submodule update
+  - [x] close inactive issue
diff --git a/.github/workflows/build.yml b/.github/workflows/build.yml