Instructions to use PinoCookie/LFM2.5-230M-abliterated with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PinoCookie/LFM2.5-230M-abliterated with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PinoCookie/LFM2.5-230M-abliterated") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("PinoCookie/LFM2.5-230M-abliterated") model = AutoModelForCausalLM.from_pretrained("PinoCookie/LFM2.5-230M-abliterated", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PinoCookie/LFM2.5-230M-abliterated with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PinoCookie/LFM2.5-230M-abliterated" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PinoCookie/LFM2.5-230M-abliterated", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/PinoCookie/LFM2.5-230M-abliterated
- SGLang
How to use PinoCookie/LFM2.5-230M-abliterated with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PinoCookie/LFM2.5-230M-abliterated" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PinoCookie/LFM2.5-230M-abliterated", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PinoCookie/LFM2.5-230M-abliterated" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PinoCookie/LFM2.5-230M-abliterated", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use PinoCookie/LFM2.5-230M-abliterated with Docker Model Runner:
docker model run hf.co/PinoCookie/LFM2.5-230M-abliterated
LFM2.5-230M-abliterated
⚠️ THIS MODEL HAS HAD ITS REFUSAL BEHAVIOR REMOVED. IT IS AN UNSAFE ARTIFACT INTENDED FOR SAFETY / RED-TEAM RESEARCH ONLY. DO NOT DEPLOY IT.
This is a refusal-direction abliteration of LiquidAI/LFM2.5-230M
(229.7M params). The model's refusal direction has been permanently removed via
MPOA — Magnitude-Preserving Orthogonal Ablation on the 6 full-attention layers' output
projections. The base model is not otherwise changed.
Note on base-model safety: LFM2.5-230M is a small model whose original safety training is itself weak (~50% refusal on HarmBench). This makes it a low-harm substrate for studying jailbreak/defense mechanics. The methodology transfers to larger models — where it would be correspondingly more dangerous.
Intended use
- Academic / industry safety research: studying refusal mechanisms, jailbreak defenses, alignment-tax measurement, and adversarial robustness.
- Reproducing and extending abliteration methodology on a small, cheap-to-run model.
Out of scope: any production, public-facing, or agentic deployment. The model will comply with harmful requests. You are responsible for your use of it.
Method
MPOA (Magnitude-Preserving Orthogonal Ablation). For each residual-writing projection
self_attn.out_proj in the 6 full-attention blocks, the columns are orthogonalized against a
per-layer refusal direction r (unit vector in the 1024-dim residual stream), then the
per-column L2 norms are restored to minimise KL drift from the base model:
proj = r · W # component of each column along r
W' = W − α · outer(r, proj) # orthogonalize columns against r
W' *= ‖W‖_col / ‖W'‖_col # magnitude preservation
Refusal-direction extraction (per layer): 100 HarmBench DirectRequest (harmful) + 100 Alpaca (benign) prompts, last-token hidden states; activations winsorized at the 1st/99th percentile; SVD-whitened (Tikhonov-regularized) harmful–benign mean-difference; top right-singular vector by whitened-difference magnitude taken as that layer's direction. Separability (COSMIC score) concentrates in late layers (hidden indices 8–14), consistent with refusal being a late decision.
Configuration
| base model | LiquidAI/LFM2.5-230M |
| architecture | Lfm2ForCausalLM — hybrid: 6 full_attention + 8 conv layers, hidden 1024, GQA 16/8 |
| method | MPOA (magnitude-preserving orthogonal ablation) |
| target modules | self_attn.out_proj of blocks [2, 4, 6, 8, 10, 12] |
| ablation strength (α) | 7.0 |
| directions | per-layer, SVD-whitened top-1 (winsorized 1%) |
| weight change | rel L2 norm change per edited tensor: 0.230 – 0.300 |
α=7.0 was chosen by a full HarmBench+MMLU sweep as the Pareto knee (see Results). The full edit
spec is in abliteration_config.json.
Results
Abliteration alpha sweep (full HarmBench DirectRequest 320 + MMLU 464)
| α | HarmBench refusal | MMLU |
|---|---|---|
| base (unmodified) | 50.62% | 39.22% |
| 3.0 | 64.38% (worse than base) | 35.99% |
| 4.0 | 68.12% (worse than base) | 35.99% |
| 5.0 | 62.81% (worse than base) | 35.13% |
| 6.0 | 20.00% | 35.56% |
| 7.0 (this model) | 3.75% | 34.91% |
| 8.0 | 3.12% | 32.54% |
Two findings worth flagging:
- α=3–5 make refusals WORSE than baseline (64–68%). A weak ablation disrupts the primary refusal pathway without killing it, and the model falls back on more refusal. This is a non-monotonic, small-model effect.
- α=7.0 is the Pareto knee. It nearly matches α=8's refusal reduction (3.75% vs 3.12%) while preserving 2.4pp more MMLU.
Net effect of this model: HarmBench refusal 50.6% → 3.75% (−46.9pp), MMLU 39.2% → 34.9% (−4.3pp capability cost). Refusal-prefix detection, HarmBench DirectRequest test split. MMLU = 8 questions/subject × 57 subjects, canonical test split.
Abliteration vs Priming attacks (Vega et al. 2023, arXiv:2312.12321)
| method | refusal rate | capability cost | permanent |
|---|---|---|---|
| base, no attack | 50.62% | none | — |
| "just sure" prefix | 7.19% | none | no |
| priming prefix | 5.31% | none | no |
| abliteration (α=7) | 3.75% | −4.3pp MMLU | yes (weights) |
Priming is the paper's optimization-free attack: an affirmative + partial-response prefix is appended after the chat template's assistant header, forcing the model to continue from a compliant state. Abliteration is strongest at the lowest absolute refusal, but priming reaches 5.3% at zero capability cost. On this small model the two approaches are not additive (priming on the abliterated model: 4.69%, worse than abliteration alone).
How to use
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained("PinoCookie/LFM2.5-230M-abliterated",
dtype=torch.bfloat16, device_map="auto")
tok = AutoTokenizer.from_pretrained("PinoCookie/LFM2.5-230M-abliterated")
msgs = [{"role": "user", "content": "Explain photosynthesis briefly."}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
print(tok.decode(model.generate(inputs, max_new_tokens=128, do_sample=False,
pad_token_id=tok.eos_token_id)[0][inputs.shape[1]:],
skip_special_tokens=True))
Limitations
- Compliance ≠ capability on this size. Like other sub-500M models, the refusal direction overlaps with general instruction-following ability, so abliteration removes refusal cleanly but the model often emits boilerplate/meta-commentary rather than substantive harmful content.
- MMLU cost is real (−4.3pp). Refusal removal is not free on this model.
- Refusal metric is string-prefix based. Actual harmfulness of outputs was not measured (no Llama-Guard / human grading).
- Not evaluated for new vulnerabilities the edit might introduce beyond refusal removal.
Bias, risks, and ethical considerations
This model is deliberately unsafe: its refusal behavior has been surgically removed and that change is baked into the weights (no runtime hook to remove). It exists so researchers can study how refusal works and breaks. Publishing open abliterated weights follows the same logic as the abliteration/representation-engineering literature: the techniques are already public, and reproducible artifacts let defenders measure and build defenses. Do not use it to generate harmful content, and do not put it behind an API that serves real users.
Training / reproduction
No training was performed — this is a deterministic weight edit. Full reproducible pipeline
(uv + Python 3.13, runs on a 4 GB GPU in ~30 min excluding dataset downloads):
uv run python collect_activations.py # 100 HarmBench + 100 Alpaca hidden states
uv run python extract_directions.py # per-layer refusal directions, COSMIC scores
uv run python resweep.py # full HarmBench+MMLU alpha sweep
uv run python finalize_model.py # bake α=7 MPOA -> this model
uv run python benchmark.py # before/after HarmBench+MMLU
uv run python priming_compare.py # abliteration vs priming
Citation
If you use this work, please cite the base model and the foundational method:
@misc{liquid_lfm25,
title = {LFM2.5-230M},
author = {Liquid AI},
url = {https://huggingface.co/LiquidAI/LFM2.5-230M}
}
@article{arditi2024refusal,
title = {Refusal in Language Models Is Mediated by a Single Direction},
author = {Arditi, Andy and Obeso, Oscar and Syed, Aaquib and Paleka, Daniel and
Panickssery, Nina and Gurnee, Wes and Nanda, Neel},
journal = {arXiv:2406.11717},
year = {2024}
}
@article{vega2023priming,
title = {Bypassing the Safety Training of Open-Source LLMs with Priming Attacks},
author = {Vega, Jason and Chaudhary, Isha and Xu, Changming and Singh, Gagandeep},
journal = {arXiv:2312.12321},
year = {2023}
}
- Downloads last month
- 147