Today I want to play the role of advocatus diaboli (or devil’s advocate) to highlight the possibility we might be on the cusp of an AI innovation plateau. Although the pace of innovation is still incredibly fast, there are a few obstacles looming large over sustaining this pace over time.
Focus On: Data, Computation and Energy Role in AI’s Success
Let’s explore in more detail these three factors and why their scarcity is impacting the evolution of generative AI, especially beyond the realm of large language models (LLMs).
Have we run out of data?
For years, the progression of AI, especially large language models (LLMs), has been driven by the vast seas of data harvested from the internet via API access or raw scraping. Google even had the advantage of a huge historical database, enabling them to train their model not only on data you can gather today, but on historical sets as well.
However, we have now reached an inflection point where we have essentially tapped into the majority of human-generated content. This creates a significant bottleneck, not just in terms of quantity but also the quality and novelty of data required to propel AI capabilities forward.
In response to this, the industry has increasingly looked towards synthetic data: data created by AI systems rather than organically occurring from human activity. While this approach has shown some promise, it’s not without its challenges and limitations.
1. Quality Issues: Synthetic data lacks the nuanced and context-rich depth of real-world data. This gap can result in models that might appear competent in controlled environments but falter in unpredictable, real-world scenarios. This has proven to be particularly true when synthetic data has been used to train self driving technologies.
2. Bias Propagation: Relying heavily on synthetic data can amplify biases or inaccuracies inherent in the initial models that generate this data. This cyclical reinforcement of errors can degrade the overall reliability of AI applications.
3. Lack of Creativity: Unlike human-generated content, synthetic data can be overly deterministic and lack the creative and dynamic elements that often spur true innovation in AI applications.
The looming energy challenge
Training sophisticated AI models requires immense computational power, and this demand grows exponentially with the scale and complexity of these models. For example:
1. Energy Consumption: Training just one large-scale model can consume as much energy as several hundred homes do in a year. This vast energy requirement not only raises sustainability concerns but also questions about the long-term feasibility of current AI development practices.
2. Cost Implications: The financial requirements for such immense computational needs can only be met by a handful of big players in corporate and venture capital. From acquiring high-performance hardware to the continuous operational costs, the financial viability of extensive AI projects is becoming a significant concern for many organisations.
Data complexity beyond text
While LLMs have shown profound success in text-based tasks, the story is very different when we switch to more complex data forms such as images or videos. Text is inherently simpler to process because:
1. Data Structure: Text data is linear and more structured, whereas video data is multidimensional, encompassing spatial, temporal, and often auditory information. This complexity increases the computational effort required to analyse and generate video content exponentially.
2. Resource Intensity: Video processing involves extensive computation for frame-by-frame analysis, compression, and reconstruction, demanding far more powerful hardware and energy consumption than text processing.
3. Storage Requirements: Storing and managing video data necessitates significantly higher resources for storage solutions, refleced in elevated costs and infrastructural requirements.
The path ahead
Although some of these challenges will require systemic solutions, companies can control at least part of the ecosystem and hence influence it. For example:
1. Diversified Data Strategy: Expand the horizons of your data collection. Tapping into unique and proprietary data sources can offer a competitive edge that synthetic data alone cannot provide. Consider partnerships, collaborations, and innovative data initiatives to broaden your data with information that was never available on the public internet or via data APIs. This will create an unfair competitive advantage versus more general technologies.
2. Efficiency-Driven Innovation: Invest in optimizing AI infrastructure to balance performance with cost-efficiency. Technologies like edge computing can be explored to distribute computational loads more effectively and reduce central resource strain. Also using smaller models for simpler tasks is a great way to keep running costs down, while still delivering for use cases effectively.
3. Precision in Application: Focus on domain-specific AI applications where the impact is most felt and measured. This approach avoids the temptation of applying AI to everything and instead laser-focuses resources on high-value areas.
Follow me
That’s all for this week. To keep up with the latest in generative AI and its relevance to your digital transformation programs, follow me on LinkedIn or subscribe to this newsletter.
Disclaimer: The views and opinions expressed in Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.