Instructions to use erax-ai/EraX-GEO-V1.5-Spectrum with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use erax-ai/EraX-GEO-V1.5-Spectrum with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="erax-ai/EraX-GEO-V1.5-Spectrum") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("erax-ai/EraX-GEO-V1.5-Spectrum") model = AutoModelForCausalLM.from_pretrained("erax-ai/EraX-GEO-V1.5-Spectrum", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use erax-ai/EraX-GEO-V1.5-Spectrum with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "erax-ai/EraX-GEO-V1.5-Spectrum" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "erax-ai/EraX-GEO-V1.5-Spectrum", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/erax-ai/EraX-GEO-V1.5-Spectrum
- SGLang
How to use erax-ai/EraX-GEO-V1.5-Spectrum with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "erax-ai/EraX-GEO-V1.5-Spectrum" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "erax-ai/EraX-GEO-V1.5-Spectrum", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "erax-ai/EraX-GEO-V1.5-Spectrum" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "erax-ai/EraX-GEO-V1.5-Spectrum", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use erax-ai/EraX-GEO-V1.5-Spectrum with Docker Model Runner:
docker model run hf.co/erax-ai/EraX-GEO-V1.5-Spectrum
YAML Metadata Warning:The pipeline tag "text2text-generation" is not in the official list: text-classification, token-classification, table-question-answering, question-answering, zero-shot-classification, translation, summarization, feature-extraction, text-generation, fill-mask, sentence-similarity, text-to-speech, text-to-audio, automatic-speech-recognition, audio-to-audio, audio-classification, audio-text-to-text, voice-activity-detection, depth-estimation, image-classification, object-detection, image-segmentation, text-to-image, image-to-text, image-to-image, image-to-video, unconditional-image-generation, video-classification, reinforcement-learning, robotics, tabular-classification, tabular-regression, tabular-to-text, table-to-text, multiple-choice, text-ranking, text-retrieval, time-series-forecasting, text-to-video, image-text-to-text, image-text-to-image, image-text-to-video, visual-question-answering, document-question-answering, zero-shot-image-classification, graph-ml, mask-generation, zero-shot-object-detection, text-to-3d, image-to-3d, image-feature-extraction, video-text-to-text, keypoint-detection, visual-document-retrieval, any-to-any, video-to-video, other
EraX-GEO-V1.5-Spectrum: Standardizing Messy Vietnamese Addresses🌍 🇻🇳
Tired of "un-recognizeable" Vietnamese addresses? EraX is here to help!
EraX-GEO-V1.5-Spectrum is an open-weight model designed to standardize messy, abbreviated, and misspelled Vietnamese addresses. Built by the EraX team, this model helps businesses clean up customer and supplier address data, turning chaotic input into clean, standardized output.
We know, we know… address standards are constantly changing (yes, another re-training is coming!). But for now, EraX-GEO-V1.5-Spectrum can tackle the most challenging real-world addresses, bringing order to the chaos.
What Does it Do?
EraX-GEO-V1.5-Spectrum takes a "dirty" Vietnamese address as input and returns a standardized Vietnamese address and its English translation. It handles:
- Abbreviations
- Misspellings
- Inconsistent formatting
Example Transformations
Here are a few examples of the magic EraX can perform:
Input (Dirty Address):
400/7 le van sy, P14, q3 hcm
Output:
{
"vn_address": "400/7 đường Lê Văn Sỹ, Phường 14, Quận 3, Thành phố Hồ Chí Minh",
"en_address": "400/7 Le Van Sy street, Ward 14, District 3, Ho Chi Minh City"
}
Input (Dirty Address):
1419/118/6C Le Van Luong, Xa Phuoc Kieng
Output:
{
"vn_address": "1419/118/6C đường Lê Văn Lương, Xã Phước Kiển, Huyện Nhà Bè, Thành phố Hồ Chí Minh",
"en_address": "1419/118/6C Le Van Luong street, Phuoc Kien Commune, Nha Be District, Ho Chi Minh City"
}
Input (Dirty Address):
khu nhà xưởng 04, lô đất 10.7,KCN Texhong hải hà,TT quảng hà,H.hải hà
Output:
{
"vn_address": "Khu nhà xưởng 04, Lô đất 10.7, Khu công nghiệp Texhong Hải Hà, Thị trấn Quảng Hà, Huyện Hải Hà, Tỉnh Quảng Ninh",
"en_address": "Factory area 04, Land lot 10.7, Texhong Hai Ha Industrial Zone, Quang Ha Town, Hai Ha District, Quang Ninh Province"
}
Input (Dirty Address):
Nhà máy thức ăn chăn nuôi tại khu Công Nghiệp Phố Nối A, Xã Lac Hong
Output:
{
"vn_address": "Nhà máy thức ăn chăn nuôi tại Khu Công nghiệp Phố Nối A, Xã Lạc Hồng, Huyện Văn Lâm, Tỉnh Hưng Yên",
"en_address": "Feed mill at Pho Noi A Industrial Park, Lac Hong Commune, Van Lam District, Hung Yen Province"
}
Input (Dirty Address):
Tôi sống quanh bờ kè Nhiêu Lộc, địa chỉ là 17-19-21 Nguyen Văn Trỗi, P.14 Q. Phú Nhuận, TpHCM
Output:
{
"vn_address": "17-19-21 đường Nguyễn Văn Trỗi, Phường 14, Quận Phú Nhuận, Thành phố Hồ Chí Minh",
"en_address": "17-19-21 Nguyen Van Troi street, Ward 14, Phu Nhuan District, Ho Chi Minh City"
}
Input (Dirty Address):
Nghĩa Thành, Châu Đức, BRVT
Output:
{
"vn_address": "Xã Nghĩa Thành, Huyện Châu Đức, Tỉnh Bà Rịa - Vũng Tàu",
"en_address": "Nghia Thanh Commune, Chau Duc District, Ba Ria - Vung Tau Province"
}
Input (Dirty Address):
10/5, Kp 2, Tân Mai, Biên Hòa
Output:
{
"vn_address": "10/5, Khu phố 2, Phường Tân Mai, Thành phố Biên Hòa, Tỉnh Đồng Nai",
"en_address": "10/5, Quarter 2, Tan Mai Ward, Bien Hoa City, Dong Nai Province"
}
How to Use
This model is designed to be used with vLLM for accelerated inference. Here's a quick start guide:
1. Install Dependencies:
pip install transformers vllm json_repair
2. Load the Model and Run Inference:
from vllm import LLM, SamplingParams
import json_repair
params = SamplingParams(temperature=0)
model = LLM(
model="erax-ai/EraX-GEO-V1.5-Spectrum")
prompt_quick = """
Bạn là một trợ lý AI đóng vai chuyên gia làm sạch địa chỉ.
Bạn được cung cấp địa chỉ không chuẩn.
# Tham khảo Danh sách 63 các tỉnh, thành phố ở Việt Nam:
Hà Nội, Hà Giang, Cao Bằng, Bắc Kạn, Tuyên Quang, Lào Cai, Điện Biên, Lai Châu, Sơn La, Yên Bái, Hoà Bình, Thái Nguyên, Lạng Sơn, Quảng Ninh, Bắc Giang, Phú Thọ, Vĩnh Phúc, Bắc Ninh, Hải Dương, Hải Phòng, Hưng Yên, Thái Bình, Hà Nam, Nam Định, Ninh Bình, Thanh Hóa, Nghệ An, Hà Tĩnh, Quảng Bình, Quảng Trị, Huế, Đà Nẵng, Quảng Nam, Quảng Ngãi, Bình Định, Phú Yên, Khánh Hòa, Ninh Thuận, Bình Thuận, Kon Tum, Gia Lai, Đắk Lắk, Đắk Nông, Lâm Đồng, Bình Phước, Tây Ninh, Bình Dương, Đồng Nai, Bà Rịa - Vũng Tàu, Hồ Chí Minh, Long An, Tiền Giang, Bến Tre, Trà Vinh, Vĩnh Long, Đồng Tháp, An Giang, Kiên Giang, Cần Thơ, Hậu Giang, Sóc Trăng, Bạc Liêu, Cà Mau
Nhiệm vụ của bạn là hãy làm sạch và trả về địa chỉ hoàn chỉnh từ địa chỉ được cung cấp sau đây.
Giữ nguyên vẹn địa chỉ, tên đường, khu phố, tổ, ấp, thôn, xóm, nhà, khu công nghiệp (hay KCN), chung cư, căn hộ, tầng lầu, tổ dân phố (TDP), bản, làng, quốc lộ, buôn làng, cụm công nghiệp, tên khu công nghiệo, tên nhà máy, xí nghiệp, kho, lô, khóm...
Tham khảo các cụm từ viết tắt: KCN - Khu Công nghiệp, KĐT - Khu đô thị, TDP - Tổ dân phố, KP - Khu phố, (CC hay C/c) - Chung cư
Lưu ý phải giữ nguyên vẹn thông tin của địa chỉ như
- số nhà, tên khu phố, tên tổ, ấp, thôn, xóm,
- tên khu công nghiệp (hay KCN),
- tên chung cư, lầu, căn hộ,
- tổ dân phố (TDP)
- bản, làng, quốc lộ, buôn làng
- tên cụm công nghiệp, tên nhà máy, xí nghiệp, kho, lô, khóm.
# Yêu cầu:
- Yêu cầu trả về tỉnh, thành phố là một trong 63 tỉnh, thành phố ở Việt Nam được cung cấp ở trên, tên chính xác quận huyện thị xã và xã phường hay thị trấn có trong địa chỉ hay trong bộ nhớ LLM.
- Nếu không thể tìm ra thông tin thì trả về "Không xác định".
- Lưu ý địa chỉ được cung cấp có thể không có dấu, hoặc viết dính liền vào nhau hoặc viết bằng tiếng Anh, phải gắn dấu phù hợp tên Tỉnh Thành phố Việt Nam hoặc dịch ra tiếng Việt
- Hãy lục tìm trong bộ nhớ của bạn về tên Tỉnh Thành phố Việt Nam, liên hệ với tên Quận Huyện Thị Xã, Xã Phường Thị Trấn và tên đường, xa lộ, cao tốc liên quan
- Hãy dùng phương pháp chain of though step-by-step để phân tích địa chỉ bẩn ban đầu, và địa chỉ phỏng đoán trước đó để suy luận ra địa chỉ đúng.
- Cuối cùng trả về json sau:
{"vn_address": "<str địa chỉ được làm sạch bằng tiếng Việt>",
"en_address": "<str địa chỉ được làm sạch và dịch sang tiếng Anh>"}
# Địa chỉ bẩn: {bad_address}
# Output:
"""
bad_address = "400/7 le van sy, P14, q3 hcm" # Replace with your address
prompt_in = prompt_quick.replace("{bad_address}", bad_address)
content = [
{
"role": "user",
"content": prompt_in
}
]
outputs = model.chat(content, sampling_params=params)
result = json_repair.loads(outputs[0].outputs[0].text)
print(result)
3. Model Prompt:
The prompt included to the instruction of how the address will be standardized. Keep the original prompt to obtain the performance.
Limitations
While powerful, EraX-GEO-V1.5-Spectrum may struggle with:
- Extremely dirty addresses ("chị nhận không ra" level).
- Addresses with very poor spelling or highly localized abbreviations.
Contributing
We encourage you to use, test, and provide feedback on EraX-GEO-V1.5-Spectrum!
License:
- LLaMA.3.1 8B is under llama3.1 license
Citation 📝
If you find our project useful, we would appreciate it if you could star our repository and cite our work as follows:
@article{title={erax-ai/EraX-GEO-V1.5-Spectrum: Standardizing Messy Vietnamese Addresses},
author={Nguyễn Anh Nguyên - Nguyễn T. Vy - Phạm Đình Thục},
organization={EraX},
year={2025},
url={https://huggingface.co/erax-ai/EraX-GEO-V1.5-Spectrum}
}
The EraX Team
- Downloads last month
- -
Model tree for erax-ai/EraX-GEO-V1.5-Spectrum
Base model
meta-llama/Llama-3.1-8B