started · updated
AI models show varying performance in business logic and creativity tests
Recent studies and tests have highlighted the varying capabilities and cost-efficiencies of large language models in business and creative tasks.
In a business analysis test involving a fictional agency problem, six different AI models were tasked with choosing between raising prices or investing in marketing. All six models reached the same strategic conclusion, and 17 out of 18 memos matched the correct answer key. The study noted that while more expensive models exist, lower-cost options were capable of handling the specific task, though the median cost for 1,000 such analyses varied significantly between the cheapest and most expensive models.
Separately, a large-scale study published in Scientific Reports compared AI creativity to human performance using the Divergent Association Task (DAT). The research, which tested language models against a dataset of 100,000 people, found that while GPT-4’s average score exceeded the average human score, the most creative segments of the human population still outperformed all tested models. The results suggest a strong machine middle ground, but a significant gap remains between AI and the top 10% of human creative thinkers.