Every day, we read fresh news on how AI companies are collecting vast amounts of data in a minefiled of legal uncertainties. Yet, as we quickly progress through new phases of AI development, the real story is evolving beyond this controversy. In fact, the future of AI relies more on precision and expertise than sheer volume. This week we discuss why this is the case and what to do next.
Focus On: The Shift from Quantity to Quality in AI Learning
In recent years, debates about AI’s reliance on massive datasets garnered through extensive web scraping have sparked ethical discussions. Critics highlight potential copyright breaches, while many social brands, such as Reddit, quickly started to make all their data available through feeds - for a fee. Yet, these datasets formed only the preliminary foundation on which AI models built their initial capabilities.
We should look at AI’s development very similar to how toddlers learn language—not by dissecting the meaning of every word but through exposure to daily usage of spoken language. Day after day, toddles link together all this signals which little by little acquire meaning. In other words, our intellectual contributions—research papers, books and articles included—played a fairly basic role in providing AI with a general understanding of language and context.
Throughout 2024 brands have quickly discovered that AI’s evolution is entering a new phase that emphasises quality over quantity. Huge datasets of unstructured data are not that valuable, the same applies to synthetic data. What really is needed to train and run high quality models returning truly useful outputs are high quality datasets.
For this reason, researchers are now focussed on teaching AI to reason, make thoughtful decisions, and produce insightful results rather than just replicating data it has consumed. This new direction requires a different approach: smaller, highly targeted, ethically sourced datasets curated by experts.
Forward-thinking companies and academic institutions are already paving the way, engaging specialists to develop bespoke data sets that delve deeply and accurately into specific fields. This method not only resolves legal ambiguity but also paves the way for a leap towards AI systems capable of handling strategic analyses, ethical challenges, and innovative solutions beyond current models’ limitations.
Organisations should reassess their data strategies, prioritising the curation of meaningful information rather than solely focussing on large-scale collection. A few takeaways:
1. Embrace the Change: Acknowledge that AI development is advancing and its trajectory is evolving, what looked like a big menace just a few months ago is blowing off. Secure and protect your high quality data, but do not fear AI bots scraping your publicy available information. If anything, it will help with your GenAI-SEO.
2. Invest in Expertise: Foster partnerships with experts who can develop highly specialised datasets that enhance AI’s cognitive development. Invest in research and development to remain at the forefront of AI’s capabilities.
3. Ethical Data Strategy: Establish a robust data governance framework, prioritising ethical considerations and legal compliance to mitigate risks associated with the data lifecycle.
4. Uncover your unfair advantage: Empower your teams with a deeper understanding of AI’s evolving opportunties and brainstorm on what is going to be your unfair advantage. Is it data? Is it the technology? Something else? Once the dust settles after this experimentation phase, brands who have identified a strong value proposition are going to harvest the opportunity first.
Follow me
That’s all for this week. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.