From data localisation to AI localisation: Why inference is moving to India | Artificial Intelligence News


The location of artificial intelligence (AI) processing is becoming a new consideration for companies deploying models in India. Anthropic has brought in-country inference for Claude to India through Amazon Bedrock, allowing requests made through an India endpoint to be processed across AWS regions in Mumbai and Hyderabad. OpenAI models are also available through Amazon Bedrock with India geographic inference, with requests routed between the same two regions.

 

The shift changes the scope of the data localisation debate. Earlier, the main question for an enterprise was where its data was stored. With AI systems, another question is where the model processes that data and generates a response.

  

That distinction matters for companies handling financial, health, government or other sensitive information. It also affects latency, access to computing capacity, regulatory compliance and the ability to control how AI systems operate.

The move is visible beyond global hyperscalers. Indian provider of hyperscale data centers,Yotta has launched a sovereign agentic AI platform for Indian organisations in partnership with IBM. Meanwhile, Indian AI startup, Sarvam and Centre for Development of Advanced Computing (C-DAC) are working towards an AI stack covering computing infrastructure, models and applications.

Data storage and inference are not the same thing

Data localisation refers to where information is stored. An Indian company could, for example, store customer records in an Indian data centre while using an AI service whose servers are located elsewhere to process a prompt.

 

Inference happens when a trained AI model takes an input, performs the computation needed to interpret it and produces an output. For a generative AI application, this is the part where the model actually processes the prompt and generates the answer.

 

Keeping storage and inference in the same country reduces another point at which sensitive information can cross borders.

 

Amazon’s India geographic inference profiles for OpenAI models route requests between its Mumbai and Hyderabad regions. AWS says the requests remain within the India geography, while the two regions provide a larger pool of capacity than a single region. Its Claude deployment follows the same model. Requests can move between Mumbai and Hyderabad, but not outside India.

This does not mean that every part of an AI service is suddenly Indian. The underlying model may still have been developed elsewhere and the computing hardware may come from global suppliers.

 

The distinction is important when discussing “sovereign AI”. Local inference is one layer of sovereignty, not the whole stack.

Why inference location matters

The first issue is regulatory compliance.

 

India does not have a blanket rule requiring all personal data to be stored or processed only in the country. However, the Digital Personal Data Protection Act, 2023 gives the Centre the power to restrict transfers of personal data outside India to notified countries or territories.

 

The DPDP rules also gives organisations a framework for handling personal data, but sector-specific rules remain important.

 

Financial services are a clear example. The Reserve Bank of India requires payment system data to be stored in systems located in India. RBI does allow certain processing outside India, but the data must be brought back to India after processing, within the prescribed timeline.

 

For an AI application processing payment information, the difference between storage and processing therefore becomes important. An enterprise may meet a storage requirement while still having to examine where an external AI model processes the information.

 

Keeping inference in India can reduce that compliance burden for certain use cases.

 

The other reason is performance.

 

When an AI application sends a prompt to a model running in another geography, the request has to travel to the remote infrastructure and the response has to return. For a conventional chatbot, that may not be a major concern. It becomes more important when AI is embedded in applications that make repeated model calls or when agents take several steps to complete a task.

 

This is also relevant as companies move from isolated AI experiments to systems that operate inside business processes. An agent handling documents, customer queries or internal workflows may need to make multiple inference calls rather than generate a single response.

Why companies are doing this now

The timing is linked to the move from AI pilots to production deployments.

 

Enterprises are using AI with business documents, customer information, financial records and internal systems. The risk assessment therefore changes when an AI model is no longer being used to summarise public information but is connected to corporate systems.

 

Agentic AI adds another layer. An AI agent can access applications, retrieve information and take actions based on its instructions. Keeping only the underlying data in India does not answer every question about where those operations are executed or who controls the infrastructure.

 

IBM and Yotta’s platform reflects this shift. The companies say their platform combines IBM watsonx Orchestrate with Yotta’s Shakti Cloud and Shakti Studio, covering AI agents, compute, GPU infrastructure, model development, inference and governance within India. The platform is intended for workflows including security operations, document processing and HR automation.

 

This is a move from securing a database to controlling the environment in which an AI system operates.

India is building capacity, but sovereignty has limits

India’s sovereign AI push also has an infrastructure component.

 

According to a PIB release in August, the IndiaAI Mission has approved 237 projects for subsidised compute support, accounting for 93.18 lakh GPU hours. Fifteen compute service providers have been empanelled, while a high-performance AI compute system of about 1.1 exaFLOPS is being established at the NIC Data Centre in Delhi.

 

The same release also notes that India’s compute ecosystem currently relies on globally sourced GPUs procured through empanelled compute service providers.

 

That makes the definition of sovereign AI important. An AI workload can be sovereign in terms of where it is processed and governed without the entire technology stack being manufactured domestically.

 

Sarvam’s partnership with C-DAC takes the discussion further. The company said that the agreement brings together indigenous computing infrastructure, AI models and production-ready applications. Sarvam’s models and applications are to run on C-DAC infrastructure, with plans for joint development and demonstrations for government and strategic-sector applications.

 

The objective is therefore not only to keep data inside India. It is to reduce dependence on external infrastructure across more layers of the AI stack.

The economics of local inference

There is a cost to this shift.

 

Running inference in a specific geography requires sufficient local computing capacity. Enterprises and model providers may have to reserve or build infrastructure in that country rather than use a global pool of servers.

 

The result can be a trade-off between sovereignty and cost. A global AI provider can spread workloads across regions based on available capacity. An India-only deployment has a smaller geographic pool, even when multiple domestic regions are available.

 

At the same time, local infrastructure can make sense when regulatory requirements or business contracts already require Indian processing. The cost of moving sensitive data between jurisdictions, assessing overseas processing arrangements and managing compliance can also form part of the calculation.

From data sovereignty to control over AI

The larger change is in what companies now mean when they ask whether their AI is “local”.

 

Data stored in India does not necessarily mean the AI processing that data happens in India. Local inference closes that gap, while sovereign platforms such as the IBM-Yotta deployment attempt to extend control to compute, model deployment, agents and governance.

 

India’s approach is also moving towards a stack where models, applications and infrastructure are connected. The Sarvam-C-DAC partnership explicitly describes this as a stack extending from computing hardware to applications.

 

This makes AI localisation a broader issue than data-centre location. It covers where information is stored, where computation takes place, which infrastructure runs the model, who controls the deployment and which jurisdiction governs the resulting system.



Source link

Related posts

Around 75% of firms in India are still not using artificial intelligence, with AI adoption remaining uneven across the economy : World Bank

India’s Gen Z Faces Down The Artificial Intelligence Future

Around 75% of firms in India are still not using artificial intelligence, with AI adoption remaining uneven across the economy : World Bank