Seymour Cray Keynote Address The following text is an invited talk given at Petaflops Workshop Pasadena, CA, January 1994. Seymour Cray Cray Computer Corporation I understand we could characterize our group today as a constructive lunatic fringe group. I would like to start off presenting what I think is today's reality, but then I'll move into the lunatic area a little later. I would like to give you three impressions today. The first one is my view of where we are today in terms of scientific computer technology. The second one is, what's the rate of progress that we, Cray Computer Corporation, are making incrementally? By incrementally I mean in a few years. And thirdly, I'd like to speculate on what I would do if I were going to take a really radical approach to a revolutionary step such as we are talking about in this workshop. What I want to do is talk about the things I know myself, and I think we are representative of where other companies are as well. In order to have some real numbers to be specific, I'll talk about my own work for a few minutes. The CRAY 4 computer is a current effort and we should complete the machine this year. We should look at the number of Gigaflops -- and this is where I'd like to start -- and at what they cost in today's prices. I would like to separate the memory issue for just a moment from the processors because they are somewhat different. If I do that, then the cost per Gigaflops in a CRAY 4 is $80,000. Now, I look at the incremental progress and project it four years, and I use four years because that's the kind of step we do in building machines: two years is too fast, but four years is about right. So if I use four years as my increment of time, and I ask what do we expect to do in that time, this gives us a rate of change. I see a factor of four every four years and I have every reason to believe that in the next four years we can continue at that rate. Whether we can continue at that rate forever, I don't know, but it is a rate that has some history and some credibility. If I look forward four years, we are going to have a conventional vector machine with about $20,000 per Gigaflops, for the processor. What does it cost per Teraflops? We are talking $20,000,000. Now we have to add memory. One of the rules of thumb we have in vector processing is for every Gigaflops in processor you need a Gigaword per second bandwidth to a common memory, and this makes the memory expensive. It's the bandwidth more than memory size that actually determines the cost. The memory cost varies somewhat from a minimum of about the cost of the processors to twice the cost of the processors. So if I pick a number, in between or 1 1/2 times that for a very big system, we would find we would have a Teraflops conventional vector machine in four years for around $50,000,000. I think that's reality without any special effort apart from normal competition in the business. I'd like to look at the other end of the spectrum because I have been involved in that recently. By the other end of the spectrum, I mean a step from a cost of $80,000 or $20,000 per processor to the other extreme end, about $6 per processor. That, in fact, is another machine we are building at Cray Computer. If we look at SIMD bit-processing, that is the other end of the spectrum, so to speak. Of course, the purpose of building this is not to do Gigaflops, Teraflops, or Petaflops but to do image processing, but never mind that for a moment. I want to come up with a cost figure here. What we are building is a 2,000,000-processor SIMD machine and it will cost around $12,000,000 to build. We are planning to make a 32,000,000-processor system in four years, and that will have a Peta-operation per second. My point is, if you program bit processors to do floating point, which may not be the most efficient thing in the world, you still come up with a machine that can do around Teraflops in four years. Whether you take a very large processor or a very small processor, either way we come up with about a Teraflops and about $50,000,000 in four years. I suspect, although I don't really know, that if we try various kinds of processor speeds in between, we're going be somewhere in the same ballpark. So my conclusion is that in four years we could have a Teraflops and it ought to cost about $50,000,000 and the price ought to drop pretty fast thereafter. So, how do we get another factor of a thousand? Well if we are able to maintain our current incremental rate, it will take 20 years. Now that might be too slow: I don't know what our goals are in this exercise. I suspect it might take 20 years anyway, but if we'd like to have both belt and suspenders, we could try a revolutionary approach and so I have a favorite one that I would like to propose. It's probably different from everyone else's. I think in order to get to a Petaflops within a reasonable period of time, or 10 years, we have to somehow reduce the size of our components (see, I am really a device person) from the micron size to the nanometer size. I don't think we can really build a machine that fills room after room after room and costs an equivalent numbers of dollars. We have to make something roughly the size of present machines, but with a thousand times the components. And, if I understand my physics right, that means we need to be in the nanometer range instead of the micrometer range. Well, that's hard, but there are a lot of exciting things happening in the nanometer-size range right now. During the past year, I have read a number of articles that make my jaw drop. They aren't from our community. They are from the molecular biology community, and I can imagine two ways of riding the coat tails of a much bigger revolution than we have. One way would be to attempt to make computing elements out of biological devices. Now, I'm not very comfortable with that because I am one and I feel threatened. I prefer the second course, which is to use biological devices to manufacture non-biological devices: to manufacture only the things that are more familiar to us and are more stable, in the sense that we understand them better. What evidence do we have that this is possible? Two areas really have impressed me, again, almost all during the past year. The rates of understanding in the nanometer world are just astounding. I don't know how many of you are following this area, but I have been attempting to read abstracts of papers, and some of them are just mind-boggling. Let me just digress for a moment with the understanding of the nanometer world as I perceive it with my superficial knowledge from reading abstracts. First, I once thought of a cell as sort of a chemical engineer's creation. It was a bag filled with fluid, mostly water, with proteins floating around inside doing goodness knows what.' Well, my perception in the past year has certainly changed because I understand now they're not full of water at all. And if we look inside, as we are beginning to do with tools that are equally mind-boggling, we see that we have a whole lot of protein factories scattered around, hundreds and thousands of them in a single cell, with a smaller number of power plants scattered around, and a transportation system that interconnects all of these things with railroad tracks. Now, in case any of you think I'm on drugs, I brought some documentation. You can read these government-sponsored reports, which you have to believe are real, because it's our tax dollars that pay for this. But I'm coming to the part that's most interesting to me. Using laser tweezers, which has been the big breakthrough in seeing what's going on in the nanometer world, human researchers have been able to take a section of the railroad track of the cell, put it on a glass slide, and lo and behold there's a train running on it with a locomotive and four cars. We can measure the speed and we did. The track is not smooth. It has indents in it every 8 nanometers. It's a cog railroad. When we measure the locomotive speed, we see it isn't smooth. The locomotive moves in little 8 nanometer jerks. When it does, it burns one unit of power from the power plant, which is an ATP molecule. So it burns one molecule and it moves one step. Well, how fast does it do this? It does it every few milliseconds. In other words, the locomotive moves many times its own length in a second. This is a fast locomotive. I am obviously impressed with the mechanical nature of what we are learning about in the large molecule world. What evidence is there that we could get anything to make a non-biological device? Or, to come right to the point, how do we try in bacteria to make transistors? Well I don't know how to do that right now, but last spring there was a very interesting experiment in cell replicating of copper wire. It's a nano-tube built with a whole row of copper atoms. The purpose of the experiment was not to make a computer, it was to penetrate the wall of the cell and measure their potentials inside without upsetting the cell's activity. These people are in a different area of concern here. But, if indeed we can make copper wire that grows itself, and this copper wire was three nanometers in diameter insulated and if we can do that today, isn't it conceivable that we can create bacteria that make something more complicated tomorrow? So, what course of action might we take to explore nanometer devices that are self-replicating? It seems to me we have to have some cross- fertilization among government agencies here. There are people doing very worthwhile research in the sense of finding the causes and cures for diseases, and more power to them, keep going. But maybe we can fund some research more directed toward making non-biological devices using the same nanometer mechanisms. So, that's my radical proposal for how we might proceed. I don't really know what kind of cross-fertilization we can get in this area, or whether any of you think this is a worthwhile idea, but it's going to be interesting for me to hear your proposals on how we get a factor of a thousand in a quick period and this is just one idea. I thank you and am ready to hear your ideas.