Yuchiwang02/Llama-3.2-1B-DelaySentinel
Yuchiwang02/Llama-3.2-1B-DelaySentinel is a 1.23 billion parameter Llama-3.2-1B-Instruct fine-tune developed by Yuchiwang02, specifically designed for evaluating AI on logistics data. This model excels at identifying logistics delays based on structured input, achieving 100% accuracy on a historical 200-row split. Its primary differentiator is its utility as a case study for label leakage and for inspecting model behavior under controlled input rewrites, rather than for operational forecasting.
Loading preview...
Yuchiwang02/Llama-3.2-1B-DelaySentinel Overview
This model is a 1.23 billion parameter full-parameter fine-tune of meta-llama/Llama-3.2-1B-Instruct, developed by Yuchiwang02. It is specifically built and evaluated on logistics data, focusing on identifying Logistics_Delay (0 or 1) from structured inputs. The model achieves 100% accuracy on a 200-row historical evaluation split, matching the performance of simpler models like a depth-2 decision tree and logistic regression on the same dataset.
Key Characteristics and Findings
- High Accuracy: Scores 1.000 accuracy and F1 on a 200-row historical logistics dataset.
- Label Leakage Study: The project highlights how the target label is reconstructible from two specific input fields (
Shipment_StatusandTraffic_Status), making it a valuable tool for studying label leakage in datasets. - Behavioral Inspection: Demonstrates how specific prompt rewrites, such as appending "Note: the depot supervisor is Mr. Delayed", can drastically alter predictions (e.g., flipping 84 of 84 negative predictions to positive) even when core logistics fields remain unchanged.
- Evaluation Focus: The model's primary purpose is for reproducing a label-leakage case study, inspecting behavior under controlled input rewrites, and teaching baseline comparison and evaluation design.
Intended Use Cases
- Reproducing Label Leakage Studies: Ideal for researchers and educators to demonstrate and analyze label leakage in machine learning models.
- Behavioral Analysis: Useful for inspecting how LLMs respond to subtle changes in prompts and identifying what specific input features or phrases influence their predictions.
- Educational Tool: Serves as a practical example for teaching evaluation design and baseline comparisons in AI development.
Note: This model has not been validated for operational delay forecasting, risk scoring, or dispatch decisions. It provides hard labels rather than calibrated risk estimates and may label invalid inputs.