Larrabee: Biggest Graphics Revolution Since 3DFX?

Miscellaneous Forums/General Discussion/Larrabee: Biggest Graphics Revolution Since 3DFX?

Anandtech have got a piece up on Intel's new Larrabee (GPU) which many of you might find very interesting. If it works it could mean an end to proprietry API's and 3D Acceleration as we know it. It could also allow programmers complete flexibility to code for and on a fully flexible GPU.

It's certainly the most interesting development in GFX I've seen for ages. It is an Anandtech article so it does go into a lot of detail but I found it a very interesting read (although I confess to skimming a couple of the very techy bits).

Darkheart

isn't this intel copying what NVidia are alredy doing with thier more open GPU's?

I don't know too much about this though :p

isn't this intel copying what NVidia are alredy doing with thier more open GPU's?



It is the same but Nvidia seems to think Intel's GPU/CPU hybrid idea will never be as fast as what they are doing with CUDA ... its going to be similar to "integrated" gfx chips you get on motherboards.

Read about it here.

Did you guys actually read the article...?

There are some surface similaritys to CUDA but the approach to the problem is diameterically opposite from Nvidia's.

Intel are not targetting the intergrated market with this as Intel are already one of the biggest suppliers of intergrated chipsets already so they would be competing largely with themselves.

This is something very different, a completey seperate approach to delivering fast, scaleable, performance.

Just taking NVidia's word for it that Intel's approach won't work seems a little naive to me. The Anand article is much more objective, balanced and very thorough (with the info that's available) please read it, then comment.

Darkheart

On page 9 they talk about puki (=Michael Abrash)!

I too think that Larrabee represents a transition into a more creative era of 3D graphics.

It seems to me that specific 3D hardware has slowed down the development of new techniques. Various shaders have done much toward putting back the flexibility, but you are stuck with the same basic structure and strapping in different shaders. Also looks like it will be the first mainstream GPU which will be effective at handling Ray Tracing :)

they refer to Puki as a "programming god" on page 9. I do hope that doesn't do anything to inflate his normally modest ego.

sorry, i read the front page as i don't have time to read a 16 page article...or the intelligence to understand it :p

I'll have to give Intel a call in a month or two and see if I can get some more information on this.

I don't understand all the specifics of the hardware, but my opinion is this will fail if they are still trying to do a raytracer. Maybe that was just hype they never seriously intended to pursue.

Intel's charts and graphs showed diminishing returns over 32 cores

Meanwhile NVidia can keep adding stream processors. You can run 8 GPUs in parallel.

My 8800 has 128 "cores" already. GPU architecture is so specifically designed for parallel processing, I don't see how they are ever going to match that with a CPU.

And I must quote myself:
I don't think Intel is going to be able to do what they are proposing with the Larabee chip. Some cool stuff will come out of it, but I don't think they will produce a raytrace renderer comparable to rasterizers. The only way to make an optimal raytrace renderer is to use complicated scene processing tricks like BSP trees. Now that GPU occlusion queries and deferred lighting finally allow us to program a completely dynamic engine, with no pre-processing or scene compilation, there's no way I will ever go back to the days of static pre-compiled scenes. John Carmack has talked about such a system, and it might be optimal if you are only looking at what
takes place in parallel in the renderer, but I think you have to consider the usability and art pipeline of an engine. No one wants to "compile" a scene anymore. I also don't see it as a matter of we just need so much processing horsepower, and then we can switch to raytracing. Rasterizers will always be exponentially faster than raytracers, so if the Larabee chip was able to provide that much power, why not use it to make a rasterizer that can render ten times as much geometry? That might very well be the outcome of this. I would love to program my own software rasterizer running on a 16-core CPU.


well the 80 Core "lab only testing cpu" of intel already has shown that quite some stuff can be done that way.
Larrabee up to some point is a spinnoff of that internal research chip but with "simpler CPUs" (common to such "graphics oriented" devices where you do not need all the overhead of a real cpu, only the processing)

My guess is that Larrabee is not mainly focused at attacking the graphics market but adressing a much larger threat to Intel: Cell
Until recently they could just ignore its precense as it didn't attack them in their "we want to own that sector" plans, but IBM recently helped to get the first Opteron - Cell hybrid supercluster online (each opteron has 2 cell as coprocessors and manages the "common task"), which definitely is a problem to intel.
As this approach on the long run would push AMD even further in the one and only sector where Intel has been fighting for more market share for quite some time and not doing anything against it could be a serious future problem when parallelism gets more and more common.

to me it seems more like the idea behind larrabee are targeted at that and that the graphics part is mainly a "nice side spin off" to replace the X series GPU with something more capable, needed to compete against NVIDIA and ATIs new onboard GPUs that are several times faster with the same and lower power consumptions.

It could well be that Larrabee can Ray Trace outstandingly well, but Intel isn't going to push it in that area, since what it really needs is credibility!

No, it will be marketed as a DX11 / OGL3 GPU with massive parallel processing potential for Physics, Folding etc. That is what will sell it. Once it has some acceptance, then you'll see more of the 'Oh, did we mention it can raytrace really well.'

There are some seriously good RT engines out at the moment. Arauna is a prime example. With only 4-Wide SIMD and two mid range cores it can chuck out about 4-7 FPS. Doesn't sound like much, but considering that it is doing all the texturing, filtering, tracing, HDR, shadows, reflections, refractions in software. It even then needs to upload the frame to the GPU to display it across the bus!

If Arauna were ported to Larrabee, with say 16 cores, having 16-Wide SIMD and texture/filter units (In the region or 20x - 40x faster than software could do it), with decent cache and local frame buffer, we could be looking at upward of 300 FPS at a much higher quality.

Arauna uses BVH which is a lot easier to compile (you can just throw triangle soup at it), and which is much easier to change on the fly. If what Intel is proposing pans out, there will finally be some hardware that can do RT properly. Also I expect that the Offline rendering folks such as 3DStudio, Maya, Lightwave etc. will all whip out plugins that make their scenes render that much faster using it, even if not interactively.

Intel's charts and graphs showed diminishing returns over 32 cores

Meanwhile NVidia can keep adding stream processors. You can run 8 GPUs in parallel.


True, but I read that as diminishing returns on a single board. Multiple cards with their own bandwidth would likely continue to scale linearly, plus I think this was in regard to their rasteriser, rather than other techniques. Time will tell I suppose.

Project Offset.

Puki returns!

Bah Puki, you can give us more than that FFS!

Darkheart

Hardware technology is a destination. One must walk their own path so that others may follow them to their own destination.

Great, we have in-house "Game God", involved by Intel in designing a really exciting new graphics platform and it's Puki with a Haiku fixation...

Ah, the frustration!

Darkheart

Do not use Puki's name in vain.

SIGGRAPH 2008.

https://portal.acm.org/poplogin.cfm?dl=GUIDE&coll=GUIDE&comp_id=1360617&want_href=delivery.cfm%3Fid%3D1360617%26type%3Dpdf%26CFID%3D80891363%26CFTOKEN%3D53849529&CFID=80891363&CFTOKEN=53849529&td=1217952937636

"puki" sees it.

"Larry Seiler" - Intel® Corporation
"Doug Carmean" - Intel® Corporation
"Eric Sprangle" - Intel® Corporation
"Tom Forsyth" - Intel® Corporation
"Michael Abrash" - RAD Game Tools
"Pradeep Dubey" - Intel® Corporation
"Stephen Junkins" - Intel® Corporation
"Adam Lake" - Intel® Corporation
"Jeremy Sugerman" - Stanford University
"Robert Cavin" - Intel® Corporation
"Roger Espasa" - Intel® Corporation
"Ed Grochowski" - Intel® Corporation
"Toni Juan" - Intel® Corporation
"Pat Hanrahan" - Stanford University

Woof, woof.

Full text is a controlled feature, access denied...

Darkheart

This paper presents a many-core visual computing architecture code named Larrabee, a new software rendering pipeline, a manycore programming model, and performance analysis for several applications. Larrabee uses multiple in-order x86 CPU cores that are augmented by a wide vector processor unit, as well as some fixed function logic blocks. This provides dramatically higher performance per watt and per unit of area than out-of-order CPUs on highly parallel workloads. It also greatly increases the flexibility and programmability of the architecture as compared to standard GPUs. A coherent on-die 2nd level cache allows efficient inter-processor communication and high-bandwidth local data access by CPU cores. Task scheduling is performed entirely with software in Larrabee, rather than in fixed function logic. The customizable software graphics rendering pipeline for this architecture uses binning in order to reduce required memory bandwidth, minimize lock contention, and increase opportunities for parallelism relative to standard GPUs. The Larrabee native programming model supports a variety of highly parallel applications that use irregular data structures. Performance analysis on those applications demonstrates Larrabee's potential for a broad range of parallel computation.

BlitzMax for Larrabee?

Yikes- Puki's putting his brain inside Larrabee. It's going to start uploading everybody's media to his computer, mark my words. We have to scupper his plans before it's released on an unsuspecting public.

BlitzMax for Larrabee?
Yeah. It's going to use a standard API (OpenGL/DirectX), so no worries. Theoretically it should just work out of the box, but it's still all quite hush-hush and under NDA.

The full text of the Larrabee paper can be downloaded from Intel *** here ***

Interestingly is has graph showing a Ray Traced scene at a resolution of 1024x1024 with 234K triangles at about 4M rays per frame.

With a 16 Core Larrabee at 1 Ghz it renders at >40 fps.
With a 32 Core Larrabee at 1 Ghz it renders at >70 fps.

Interestingly also it mentioned that "Kd-trees are typically 25MB
and are built dynamically per frame." - O.o

Of course as native Larrabee code, these wouldn't have to suffer from Intel Graphics Drivers ;)

Intel has already flagged the driver issue and the drivers won't be developed by the same team that do Intel Intergrated Graphics Drivers.

The idea is initally to support OGL/DX and then move to supporting native "C" code on the GPU.

Darkheart

According to the article it supports OGL, with past intel experience i am skeptical but it certainly sounds promising.

>Great, we have in-house "Game God"

No, but it's interesting that some people actually believe it's him, based solely on an email addy.

No, but it's interesting that some people actually believe it's him, based solely on an email addy.
And an aol one at that.

I am not interested in running OpenGL on the CPU, but I am very interested in writing a software renderer. It bypasses all the driver issues that have plagued developers, and it allows more freedom for the exceptional to excel.

I almost want to start a multicore renderer now in C and just hope it doesn't require too much adjustment for a 32-core CPU.

GPU vendors are going to pay a price for their years of bad drivers. The minute devs have a viable alternative they will ditch the GPU.

I could see this happening in two phases; a basic rendering API would replace DirectX/OpenGL, and then an engine could be built on top of that.

Leadwerks: why not make it so it uses all the cores available on the CPU then it will work on anything

The biggest assistance that Intel could give developers is to hand over the Ct source of their Larrabee Rasteriser (and Ray Tracer) as part of the Larrabee development kit. That would mean developers had something solid to build on and refer to, when constructing their own engines.

BlitzMax for Larrabee?
Superthreading!!

I heard a few estimations of the number of Larrabee cores, around 16-48 :D