Exploratory Bandit Experiments with “Starter Packs” in a Free-to-Play Mobile Game

Экспериментальные бандитные исследования с «стартовыми наборами» в free-to-play мобильной игре
Julian Runge, Anders Drachen, William Grosso
2024-08-05

cold start with prior offer policycontextual banditsconversion-reward optimizationonline reinforcement learning systemstarter pack assignment
This paper explores the application of bandit methods for the assignment of “starter packs” to new players in a free-to-play mobile game environment. Leveraging an online reinforcement learning system, the study aims to strategically assign starter packs to players from different country and device segments. The architecture of the online experimentation system enables real-time decision-making processes and continuous model tuning. The cold start problem is addressed by seeding the bandit with a prior offer policy informed by institutional expertise. Offline evaluation on a dataset derived from a previous AB test conducted by the company and two online experiments assess the abilities of bandit methods to assign starter packs to new players in this environment. While personalization of starter pack assignment in contextual data (country and device segments) is not achieved, a bandit with a conversion-reward changes the prior institutional policy and lowers the average effective sales price of starter packs to new players. The bandit’s policy achieves an indicative lift in per-user revenue, repeat purchasing, and player retention compared to a holdout group with a naive policy. This research contributes to the field of monetization strategies in mobile free-to-play games, emphasizing the design of systems for online personalization and revenue optimization.
1
A bandit optimized for conversion reward modified the institutional prior policy and reduced the average effective sales price of starter packs to new players.
2
An online reinforcement learning (bandit) system was implemented to assign starter packs to new players in a free-to-play mobile game, enabling real-time decision-making and continuous model tuning.
3
Offline evaluation using a previous AB test dataset and two online experiments showed that contextual personalization by country and device segments was not achieved.
4
The bandit policy produced indicative lifts in per-user revenue, repeat purchasing, and player retention versus a holdout group using a naive policy.
5
The cold-start problem was mitigated by seeding the bandit with a prior offer policy based on institutional expertise.

Assignment of starter packs to new players in a free-to-play mobile game using bandit methods and an online reinforcement learning system

Effectiveness of bandit-based assignment (including seeding with prior institutional policy, conversion-reward formulation, and real-time model tuning) on sales price, per-user revenue, repeat purchases, and player retention across country and device segments

Publication Details
Publication Date
2024-08-05
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Julian Runge
Anders Drachen
William Grosso
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%