IBOU-SEARCH/ArseneLupin-v1.1
ArseneLupin v1.1 by IBOU-SEARCH is a 4.5 billion parameter open decision model based on Qwen/Qwen3.5-4B, optimized for classification and decision-making tasks. It processes a given state and typed questions (yes/no, choice, score) to return a probability distribution for each option in a single forward pass, without generating text. This model excels at tasks requiring precise, probabilistic answers rather than generative responses, offering high accuracy on various proprietary and public benchmarks.
Loading preview...
ArseneLupin v1.1: A Specialized Decision Model
ArseneLupin v1.1 is a 4.5 billion parameter open decision model developed by IBOU-SEARCH, built upon the Qwen/Qwen3.5-4B architecture. Unlike traditional generative LLMs, this model is designed to take a 'state' (e.g., a document, ticket, or JSON record) and specific, typed questions, returning a probability distribution for each potential answer. This allows for highly efficient and precise classification and decision-making without generating free-form text.
Key Capabilities
- Typed Questions: Supports
noul(yes/no),choice(one option from several), andscore(ordered level) question types. - Probabilistic Outputs: Provides a confidence score for each answer option, enabling nuanced decision-making.
- Efficient Processing: Designed for a single forward pass per option, making it fast for scoring multiple choices.
- Performance: Achieves strong results on various benchmarks, including proprietary tasks and public datasets like Banking77 and CLINC150 intents. Notably, it outperforms Jev 1.13 on transfer to untrained decisions (92.0% vs 85.0%) and hate speech detection (B/2 0.0658 vs 0.2016).
- Calibration: Features optional per-workflow temperature calibration to fine-tune probability accuracy.
- Speed Optimization: When scoring multiple options, it processes shared text once, significantly reducing inference time per option (e.g., 19 ms per option vs. 79 ms in default mode).
- Flexible Deployment: Available in a complete bfloat16 version (9.3 GB) and an 8-bit GGUF Q8_0 version (4.5 GB) for
llama.cpp.
Good For
- Automated ticket routing and classification.
- Incident severity assessment.
- Intent recognition in customer service applications.
- Content moderation (e.g., hate speech detection).
- Any use case requiring precise, probabilistic classification or decision-making from structured or unstructured input, where text generation is not needed.