Implement CFG for ExLlama_HF (#3666)

2023-08-24 16:27:36 -03:00 · 2023-08-24 16:27:36 -03:00 · d6934bc7bc
commit d6934bc7bc
parent 2b675533f7
8 changed files with 122 additions and 26 deletions
--- a/README.md
+++ b/README.md
@ -304,6 +304,7 @@ Optionally, you can use the following command-line flags:
 |------------------|-------------|
 |`--gpu-split`     | Comma-separated list of VRAM (in GB) to use per GPU device for model layers, e.g. `20,7,7` |
 |`--max_seq_len MAX_SEQ_LEN`           | Maximum sequence length. |
+|`--cfg-cache`                         | ExLlama_HF: Create an additional cache for CFG negative prompts. Necessary to use CFG with that loader, but not necessary for CFG with base ExLlama. |

 #### GPTQ-for-LLaMa