- Chip security must be designed in from the start, with clear requirements that connect architecture, verification, software, manufacturing, and field deployment.
- AI is expanding both the scale and accessibility of attacks, making hardware-level vulnerabilities harder to anticipate, patch, and contain.
- Regulations, customer expectations, and lifecycle risks are turning semiconductor security from a technical feature into a business-critical requirement.
Fig. 1: L-R: Synaptics’ Arora, Synopsys’ Hinkel, Arteris’ Siwinski, Cadence’s Vunnam, Keysight’s Petr, Rambus’ Best, and Siemens’ Giles. SE: What are the most urgent security issues for chip and system architects to contend with today? Siwinski: The most important thing right now, and the key thing, is to just start. There is a consideration of how this will evolve — the technologies and how you solve it — but you have to start in planning it and having the conversations right now between the security team, the designers, and the verification teams, with security as part of the shift level in the design phase, and what it actually means. It’s important to get going, even if you don’t get it right. If you’re not starting now, you’re too late. This is in contrast to what’s happening right now, where security teams have some high-level understanding of what security in the chip is, or their end product is, and they have some ways to explain it to the legal department so they can say, ‘yea.’ And the designers are just understanding that they have to make sure they don’t screw up. The verification teams are overloaded, and they’re saying, ‘As long as we get functionality correct, it’s going to be good.’ But that’s not enough. You have to design it with security assurance being a core tenet of the design process, and you have to start the journey. That is the most important thing right now. Even if the destination might be shifting — and it will be shifting as technologies evolve — starting now is the most paramount thing. Hinkel: Just start, but also understand what the starting place is, because there is a baseline that’s required for products. If you’re going to ship in certain territories in the world, if you can’t support that, then your product may not be able to be sold. Maybe you can make it work with software patches, but it’s not going to perform the way that you want, and it’s probably going to come up short in some of the security assessments. One of the things we’ve noted with customers is that it might be best to refer them to a lab first to actually go through and say, ‘What do you want to build? Do you realize that you’re going to run into these specifications and certifications and things that you’re going to need for your product to be sold the way that you want it sold?’ Get them trained up to the point where they’re not so resistant when you start talking about what it’s going to take. Different companies can offer different kinds of solutions in different bits, shapes, and sizes. That’s all good, but security can’t be glued on at the end. It has to be foundational. If you don’t have certain foundational elements, you’re going to fail some of the PPA. You’re going to fail some of the features. You’re going to fail on security evaluations, whatever it happens to be. And you don’t want to spin a chip to do that. Our broad industry view is that we can also extend that to digital modeling to show why they need to be inserted. Some of the modeling that we can do on top of our tools includes side-channel power analysis and other things that are part of our tool flow. The important thing is that once they understand how easy it is to circumvent things if they haven’t built in the right defenses, then they might take action and understand that they’re going to have to spend a little bit more on the IP. They’re going to have to spend a little bit more area. They’re going to have to spend in areas where they hadn’t considered. But if they don’t understand what their security target is in the beginning, and at least have a minimum viable baseline, they’re probably going to fail with their product in the long run. Arora: A foundation and all those things are very important, and given that the attack profile is changing, the foundation is going to change. What we are also seeing is security moving from an engineering problem to mostly a business problem now. That’s coming up in a very big way. One example is the Cyber Resilience Act. Historically, it was, ‘I want to design a secure product. The brand is okay. Prove it.’ You have to go through all the processes and procedures to make sure it’s not just about security. It’s that you can establish trust. How do you maintain trust across the lifecycle of the product? That’s one of the key aspects. And our customers, the integrators, were becoming very particular about how to handle vulnerabilities. Beyond the more foundational stuff, what we are seeing is the influx of all these regulatory acts that are pushing us to think in a very different direction. Vunnam: One focus point for me and our R&D team is basically what you want an LLM to secure. Most people are focused on LLM data privacy, but at the silicon and data center level, we believe the major threat is to integrated corruption. Your data defines what you want to execute. If an attacker can ingest a prompt or tamper with the model weights, then they can alter the execution path. And with how LLMs are designed and how we keep using the data again and again, once you have a vulnerability, do you know if the next time you use a similar type of data, you actually have a guarantee the data is secure? And how do you know your agentic workflow is secure? So that’s an interesting topic which we’ve been discussing. Petr: Since we’re talking about security for chips, and we’re also talking about security for the design stack, similar to what has been said earlier, a general mindset shift is needed. You could articulate this in the same way as we’re talking about multi-physics. Everything just becomes a bigger problem nowadays because we can address bigger problem spaces with new technologies such as LLMs. As soon as we go broader and bigger, integrations, usability, and all of this begs the question about who is driving what. That drives the requirement. The CAD team says the same. The thermal guy says the same. The analog guy says the same. The digital guy says the same. So basically, it’s a convolution of everyone wanting to own the world and drive the requirements. That’s an interesting conversation. Ultimately, you need to make that part of your specification. If it’s not specified, it’s not going to happen. On the software stack, we’re also facing things like SSDF (software security design framework). So even though it got booted down a little bit, we are now, as software vendors, required to live up to certain expectations. The European CRA is even worse because they’re telling you that if you do something wrong, your profit is on the line and you’re going to pay penalties. So the question of us sending secure software to our customers in certain domains is becoming significantly more interesting. And since we’re all now shipping some kind of co-pilot agentic framework, the question of LLMs becomes increasingly important. In some cases, we rely on frontier models, which are not owned by our companies. So you’re bringing in a new party that sits in between, securing those channels, and making sure there’s no attack possible from that side. This requires a completely new design cycle on the software side. Even if we bring our own LLMs, we bring them on-site. That is a way to mitigate those security concerns by saying, ‘We put that on different hardware. We put that in an on-prem edge deployment solution.’ With all of those things, if you talk about security, you can go very broad and everywhere. Best: Going back to what was said about just getting started, most of the people at DAC are concerned about secure services, and secure services do not work without secure software underneath it. Secure software doesn’t work unless there’s secure hardware underneath it. So coming from a hardware company, one of the things you need to get started is hardware security, because not all hardware is created equal. There’s hardware, and there’s tamper-resistant hardware. If you do not have a tamper-resistant hardware team, you’d better get one or partner with one from the start. You cannot get to the secure services without being able to chase it all the way down to a ground truth and know that it’s executing somewhere in tamper-resistant hardware. Giles: The issue is simply summarized in terms of scale. The security issues that the world faces today, and the security issues in our industry, are all about scale. There are far too few people who understand security, and far too few who understand the importance of the root of trust. And unfortunately, taking AI out of the picture, the best we’ve been able to do as an industry is say, ‘Here are all the previously exploited violations. Let’s go look for those. Let’s make sure that those aren’t happening in the current design. The problem is, it’s the stuff you don’t think of, even with specifications, even with people who really understand security. There is something out there. I don’t believe that the state space of secure and exploitable weaknesses is bounded. There’s something out there that is exploitable. Now, let’s bring AI into the picture as a human intellectual base. We’re only capable of so much AI. Unfortunately, good AI cannot help us. Bad AI is certainly going to outpace our ability to protect against attack, so we face issues from a scale perspective in terms of compute. We face it in terms of the ability to mount a defense and have the good AI, the white hat hacker AI, outpace the nefarious AI, and it’s really the knowledge base. We need scale. It’s very similar to the problems with formal verification, and not surprisingly, security needs formal verification in its exhaustive nature. There are just not enough people in the world who understand that. So, put differently, the world needs AI to help democratize security knowledge and scale it. SE: For the chip architect in charge of a design team, they need to have a plan. How are your customers handling this? Are you finding that engineering teams have an organized plan to implement security, where it is in the product spec, then traced through the design, verification, and manufacturing process, and then out into the field? Siwinski: Yes. The concepts we’re dealing with here are not new. Security is not something that suddenly, overnight, became more exciting. Yes, it got more exciting in April with Mythos, but it’s the reality of it. And that’s when the LLMs, all of a sudden, were able to not only find new weaknesses, the unknown unknowns in software, but also get to the hardware level. That’s less publicized, but those models are now effective enough that they can make their way not only through the entire software stack but also to the underlying hardware. So this is not just a software question. It’s now a silicon question, and those hardware problems are more challenging in a way because you can argue there are fewer of them because you know there are only so many ways you can do things on a chip. On the flip side, when you’re in production, when you’re deployed, this is not a question of just having a simple patch. Yes, you might get lucky. You might have something you can do a firmware thing on. Maybe you have some redundant logic. Maybe you have some extra cryptographic code that you can potentially swap. But unlike in software, where, when you find something, you hopefully can fix it right away and propagate your patch, in hardware, we’re still dealing with Spectre and Meltdown. To this day, there are variants of that still happening and being reported weekly. It’s been years, and it’s just one specific flavor that’s super well documented, and yet… Then, to the question of the customers and what’s really happening, this is not new. It comes back to, ‘Only the paranoid survive.’ Some people have been more paranoid than others. Of course, not everybody wants to be public about the paranoia and why they’re paranoid. This is one of the things that also changes, an example being what happened before with functional safety for those of you who went through the 26262 journey. Safety used to be, ‘No, you don’t talk about safety. Of course, it’s safe. You don’t want to have the conversation.’ Now it’s, ‘Of course you need to have the conversation because of regulation, because of requirements, because the legal departments will not let you take the chip, because your ASP for the product is higher if you can put the sticker on it that it’s safe. So there are commercial reasons, and the ‘however you get there’ kind of thing. It went from something that used to be a taboo topic to something that became a reality. We have to deal with it, and here are the best practices. Cybersecurity, at least in semiconductors, is going through exactly the same journey, back to the customers that are public, and it’s happening already. Arm just went through this. Arm has been doing this for years. They just went public that they’re doing this a few weeks ago, that they’re making sure that all of their processors are now built with thinking through the hardware security and what that means for the software system stack, from the full software to hardware, and what secure actually means and how to build it in. How do you build it from scratch? How do you start? How do you bring the various groups together to define the objective and, in Arm’s case, equip verification teams — who are often best suited for negative thinking — to act on it? So if you guys run positive testing versus negative testing concepts, it builds on formal. But while formal is great, it has obvious scale limitations. It can only check how you run your software loads on your hardware for the right workloads to retest them. It’s great for certain things, but it needs to be complemented by other means. There are multiple solutions to this, but basically, are there best practices based on CWEs? Based on practices and other testing and negative testing, can you orchestrate your design to have something that can be runtime executable, deployed, and scaled, and then propagated downstream? The short answer is yes. Will it evolve? It will, and yes, there are solutions out there. Hinkel: We have a range of customers. Some have lots of documentation or specifications just for security. And then, some have much more with very large teams. They probably have more people employed in their security team than some of the companies that use their products. We have others with baseline specs, which is basically a list of bullets that their marketing manager gave, and maybe cryptography shows up and things like that. But they don’t know anything about the software. They don’t know anything about it because they’re just the hardware guys. The next level of educated ones has the software guys get involved, so at least they’re covered. They know what they need to do from the software perspective, and so they drive a few more requirements. The big middle ground is one that we run into — customers that think they know, but they don’t really know, because they’re kind of self-trained, self-taught. When I came into the industry, from a security perspective, I learned CIA — confidentiality, integrity, and availability. And for many years, security was concentrated on a little bit of the I (integrity). People didn’t understand the importance of availability, which becomes incredibly important for two reasons. One, if you have autonomous systems, the functionality has to be available all the time. Otherwise, the whole system could fail, and it might end in death. It might end in something severe. It goes back to functional safety, which spends a lot of time on integrity and availability. The other thing is, when we think of hacks and attacks, most people, top of mind, are thinking about data theft or maybe taking over the device, things like that. But the most dangerous one, and probably the easiest one to get to, is denial of service. All they have to do is cut the communications or cut off some vital piece of what’s going on in a system, and then they basically cause economic damage. And if they’re smart, they’ve actually made it to where they can turn that on and off, and then say to you, ‘Here’s my crypto wallet, send me the money.’ That happens a lot. That is the number one attack out there. People think it happens because of email phishing or the like, but it happens to many devices. There are whole cities where water supplies have been held hostage, and there’s no good reporting system. You’re not mandated to report that you had the breach, which I think you should be. Now, financial and certain institutions are [mandated to report incidents], but a lot of the people who are paying these off are not mandated to report. Otherwise, people would be much more aware of the concern. Transparency in security and the security market is going to become much more critical. Arora: Starting on the availability side, at some point, the denial-of-service attacks were out of scope. It was the same for the CrowdStrike attack. But that’s not the case. They can flip the world upside down. We are seeing quite a bit of heavy lifting from nano services, and that goes to availability. That’s one of the key aspects of what has to be part of the foundation as a minimum bar — to make sure the products are available. There are two ways of doing it. One is to detect and remediate as quickly as possible, or absolutely avoid it through physical detection, and so on. But the bar for physical attacks has gone very low because with AI assistance, it’s so easy. What was previously supposed to be a million-dollar attack is no longer the case. It’s a much cheaper attack now. You can go to a third-party lab and get part of the service. It’s no longer a nation-state issue. The foundation has to have that element so that your architecture is resilient. With a more complex SoC, you’ve got to really understand the SoC and the attack profile very well to protect it. For example, a more complex SoC could have like 200 power nodes. The chip is constantly being beaten down, going in and out of various domains. How do you ensure that the root of trust, or whatever security processor, can really crack? It has to be visible and transparent in a way that it knows what’s going on. I also see security evolving from, let’s say, permission: you can deny this, deny that, to a more contextual view. ‘Why is my NPU accessing Bluetooth stack memory?’ That has to be more adaptive and resilient, and the foundational element needs to take care of that. With more attacks, easier attacks, you have to have a foundation that takes the trust — not just the boot side, also runtime — all the way through the lifecycle. The post The Next Big Chip Failure May Be A Security One appeared first on Semiconductor Engineering. Source: https://semiengineering.com/the-next-bi ... urity-one/