Text Generation
Transformers
Safetensors
gemma2
creative-writing
gutenberg
conversational
text-generation-inference
Instructions to use sam-paech/Quill-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sam-paech/Quill-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="sam-paech/Quill-v1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("sam-paech/Quill-v1") model = AutoModelForCausalLM.from_pretrained("sam-paech/Quill-v1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use sam-paech/Quill-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "sam-paech/Quill-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sam-paech/Quill-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/sam-paech/Quill-v1
- SGLang
How to use sam-paech/Quill-v1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "sam-paech/Quill-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sam-paech/Quill-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "sam-paech/Quill-v1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sam-paech/Quill-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use sam-paech/Quill-v1 with Docker Model Runner:
docker model run hf.co/sam-paech/Quill-v1
Update README.md
Browse files
README.md
CHANGED
|
@@ -19,12 +19,14 @@ model-index:
|
|
| 19 |
|
| 20 |
GGUFs here: [https://huggingface.co/mradermacher/Quill-v1-GGUF](https://huggingface.co/mradermacher/Quill-v1-GGUF)
|
| 21 |
|
| 22 |
-
Quill is a
|
| 23 |
|
| 24 |
This model was trained using gemma-2-9b-it as the base. The training methods used were ORPO (gently) then SIMPO (less gently).
|
| 25 |
|
| 26 |
It scored 79.75 on the [EQ-Bench creative writing benchmark](https://eqbench.com/creative_writing.html).
|
| 27 |
|
|
|
|
|
|
|
| 28 |
[**Gutenberg3**](https://huggingface.co/datasets/sam-paech/gutenberg3-generalfiction-scifi-fantasy-romance-adventure-dpo) is a new, large dpo dataset containing extracts from 629 public domain fiction novels in the Gutenberg Library. It follows the same format as JonDurbin's original gutenberg set. It includes pairs of texts, where the chosen text is taken directly from a novel from the Gutenberg library, and the rejected text is generated by a language model based on a description of the passage. For this dataset I've used gemma-2-9b-it to generate the rejected texts, the idea being that it should more easily steer the base model away from its normal style (as compared to generating the rejected texts with random/weaker models).
|
| 29 |
|
| 30 |
# Sample Outputs
|
|
|
|
| 19 |
|
| 20 |
GGUFs here: [https://huggingface.co/mradermacher/Quill-v1-GGUF](https://huggingface.co/mradermacher/Quill-v1-GGUF)
|
| 21 |
|
| 22 |
+
Quill is a capable, humanlike writing model trained on a large dataset of late 19th and early 20th century writing from the Gutenberg Project. This model writes with a natural cadence and low gpt-slop, having inherited some human qualities from the Gutenberg3 dataset. It writes with more simple, spare prose than the typical overly-adjectived LLM writing style.
|
| 23 |
|
| 24 |
This model was trained using gemma-2-9b-it as the base. The training methods used were ORPO (gently) then SIMPO (less gently).
|
| 25 |
|
| 26 |
It scored 79.75 on the [EQ-Bench creative writing benchmark](https://eqbench.com/creative_writing.html).
|
| 27 |
|
| 28 |
+
**Instruct Template:** Gemma
|
| 29 |
+
|
| 30 |
[**Gutenberg3**](https://huggingface.co/datasets/sam-paech/gutenberg3-generalfiction-scifi-fantasy-romance-adventure-dpo) is a new, large dpo dataset containing extracts from 629 public domain fiction novels in the Gutenberg Library. It follows the same format as JonDurbin's original gutenberg set. It includes pairs of texts, where the chosen text is taken directly from a novel from the Gutenberg library, and the rejected text is generated by a language model based on a description of the passage. For this dataset I've used gemma-2-9b-it to generate the rejected texts, the idea being that it should more easily steer the base model away from its normal style (as compared to generating the rejected texts with random/weaker models).
|
| 31 |
|
| 32 |
# Sample Outputs
|