shahidul034
/

KUETLLM_zephyr

Generated from Trainer

Model card Files Files and versions

shahidul034 commited on Dec 18, 2023

Commit

9bc7370

·

1 Parent(s): 9e77dc2

Update README.md

Files changed (1) hide show

README.md +43 -6

README.md CHANGED Viewed

@@ -7,7 +7,8 @@ model-index:
 - name: KUETLLM_zephyr
   results: []
 ---
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
 should probably proofread and complete it, then remove this comment. -->
@@ -17,11 +18,29 @@ This model is a fine-tuned version of [TheBloke/zephyr-7B-beta-GPTQ](https://hug
 ## Model description
-More information needed
-## Intended uses & limitations
-More information needed
 ## Training and evaluation data
@@ -43,7 +62,25 @@ The following hyperparameters were used during training:
 - num_epochs: 2
 - mixed_precision_training: Native AMP
-### Training results

 - name: KUETLLM_zephyr
   results: []
 ---
+KUETLLM is a [zephyr7b-beta](https://huggingface.co/HuggingFaceH4/zephyr-7b-beta) finetune, using a dataset with prompts and answers about Khulna University of Engineering and Technology.
+It was loaded in 8 bit quantization using [bitsandbytes](https://github.com/TimDettmers/bitsandbytes). [LORA](https://huggingface.co/docs/diffusers/main/en/training/lora) was used to finetune an adapter, which was leter merged with the base unquantized model.
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
 should probably proofread and complete it, then remove this comment. -->
 ## Model description
+Below is the training configuarations for the finetuning process:
+```
+LoraConfig:
+r=16,
+lora_alpha=16,
+target_modules=["q_proj", "v_proj","k_proj","o_proj","gate_proj","up_proj","down_proj"],
+lora_dropout=0.05,
+bias="none",
+task_type="CAUSAL_LM"
+```
+```
+TrainingArguments:
+per_device_train_batch_size=12,
+gradient_accumulation_steps=1,
+optim='paged_adamw_8bit',
+learning_rate=5e-06 ,
+fp16=True,
+logging_steps=10,
+num_train_epochs = 1,
+output_dir=zephyr_lora_output,
+remove_unused_columns=False,
+```
 ## Training and evaluation data
 - num_epochs: 2
 - mixed_precision_training: Native AMP
+### Inference
+```
+def process_data_sample(example):
+    processed_example = "<|system|>\nYou are a KUET authority managed chatbot, help users by answering their queries about KUET.\n<|user|>\n" + example + "\n<|assistant|>\n"
+    return processed_example
+inp_str = process_data_sample("Tell me about KUET.")
+inputs = tokenizer(inp_str, return_tensors="pt")
+generation_config = GenerationConfig(
+    do_sample=True,
+    top_k=1,
+    temperature=0.1,
+    max_new_tokens=256,
+    pad_token_id=tokenizer.eos_token_id
+)
+outputs = model.generate(**inputs, generation_config=generation_config)
+print(tokenizer.decode(outputs[0], skip_special_tokens=True))
+```