Implement a demo HF wrapper for exllama to utilize existing HF transformers decoding. (#2777)

2023-06-22 02:31:42 +08:00 · 2023-06-22 02:31:42 +08:00 · 580c1ee748
commit 580c1ee748
parent a06acd6d09
7 changed files with 101 additions and 6 deletions
--- a/README.md
+++ b/README.md
@ -212,7 +212,7 @@ Optionally, you can use the following command-line flags:

 | Flag                                       | Description |
 |--------------------------------------------|-------------|
-| `--loader LOADER`                          | Choose the model loader manually, otherwise, it will get autodetected. Valid options: transformers, autogptq, gptq-for-llama, exllama, llamacpp, rwkv, flexgen |
+| `--loader LOADER`                          | Choose the model loader manually, otherwise, it will get autodetected. Valid options: transformers, autogptq, gptq-for-llama, exllama, exllama_hf, llamacpp, rwkv, flexgen |

 #### Accelerate/transformers