An analysis of machine learning approaches for enhancing decision-making in complex discrete choice tasks
The study evaluates four ML models (multinomial logistic regression, GAM, twinned neural network, Gaussian process) against five behavioral choice rules in discrete choice modeling Semi-parametric and non-parametric models consistently outperform parametric models across all choice rules and experimental conditions Model performance improves by 6-96% with more training choice sets and 0-55% with higher choice rule determinism In a real-world energy policy case study, the Twinned Neural Network (
Analysis
TL;DR
- The study evaluates four ML models (multinomial logistic regression, GAM, twinned neural network, Gaussian process) against five behavioral choice rules in discrete choice modeling
- Semi-parametric and non-parametric models consistently outperform parametric models across all choice rules and experimental conditions
- Model performance improves by 6-96% with more training choice sets and 0-55% with higher choice rule determinism
- In a real-world energy policy case study, the Twinned Neural Network (TNN) achieved the best fit with a BIC of 13.351
- The research demonstrates that model selection should be driven by the specific choice task context rather than a one-size-fits-all approach
Why It Matters
This work bridges the gap between traditional parametric discrete choice modeling in policy-making and modern machine learning approaches, offering practitioners data-driven alternatives that better capture individual heterogeneity. For AI researchers and policy analysts, it provides empirical guidance on when and which ML models are most effective for preference elicitation tasks, directly impacting how computational tools can support evidence-based policy decisions.
Technical Details
- Models evaluated: Multinomial Logistic Regression (parametric), Generalized Additive Model (semi-parametric), Twinned Neural Network (non-parametric), and Gaussian Process (non-parametric)
- Choice rules tested: Linear strong utility, monotonic strong utility, ideal point, lexicographic semiorder, and multiattribute linear ballistic accumulator — all established in behavioral and social sciences
- Monte Carlo experimental design: Systematically varied three dimensions — number of attributes in choice alternatives, number of training choice sets, and choice rule determinism — to assess robustness under realistic elicitation challenges
- Real-world validation: Energy policy preference data case study where TNN outperformed all other models with a BIC of 13.351
- Key finding: Non-parametric and semi-parametric approaches show superior generalization, particularly under high attribute complexity and individual heterogeneity
Industry Insight
- Policy-making organizations should transition from purely parametric discrete choice models to semi-parametric/non-parametric ML approaches, especially when dealing with complex, high-dimensional choice environments where individual heterogeneity is significant
- The 6-96% performance gains from increased training data suggest that investing in richer preference elicitation datasets yields disproportionately large returns — organizations should prioritize data collection volume and quality
- Model selection should be context-driven: simpler parametric models may suffice for highly deterministic, low-attribute tasks, while TNNs and Gaussian processes are better suited for complex policy scenarios requiring nuanced individual-level preference estimation
Disclaimer: The above content is generated by AI and is for reference only.