{"id":14600,"date":"2010-01-17T09:45:15","date_gmt":"2010-01-17T09:45:15","guid":{"rendered":"http:\/\/alienbabeltech.com\/main\/?p=14600"},"modified":"2010-05-02T21:26:05","modified_gmt":"2010-05-02T20:26:05","slug":"nvidias-directx-11-architecture-gf100-fermi-in-detail","status":"publish","type":"post","link":"http:\/\/alienbabeltech.com\/main\/nvidias-directx-11-architecture-gf100-fermi-in-detail\/","title":{"rendered":"NVIDIA&#8217;s DirectX 11 Architecture: GF100 (Fermi) In Detail"},"content":{"rendered":"<p align=\"center\"><script type=\"text\/javascript\">\/\/ <![CDATA[\n  google_ad_client = \"pub-7021221536180758\"; \/* 468x60, created 11\/11\/09 *\/ google_ad_slot = \"1949837292\"; google_ad_width = 468; google_ad_height = 60;\n\/\/ ]]><\/script><br \/>\n<script src=\"http:\/\/pagead2.googlesyndication.com\/pagead\/show_ads.js\" type=\"text\/javascript\">\n<\/script><\/p>\n<p>Article written by <span style=\"color: #00ff00;\"><strong>Mark Poppin<\/strong><\/span> and <span style=\"color: #00ff00;\"><strong>BFG10K<\/strong><\/span>, AlienBabelTech Senior Editors.<\/p>\n<p><!--pagetitle:Introduction--><strong><span style=\"color: #99ccff;\"><span style=\"text-decoration: underline;\">Introduction<\/span><\/span><\/strong><\/p>\n<p>At their Graphics Technology Conference (GTC) last September 30th, NVIDIA announced their next-generation graphics architecture, codenamed Fermi. We reported on it for you <a href=\"http:\/\/alienbabeltech.com\/main\/?p=11661\">here<\/a>, <a href=\"http:\/\/alienbabeltech.com\/main\/?p=11825\">here<\/a> and <a href=\"http:\/\/alienbabeltech.com\/main\/?p=11911\">here<\/a> in a three-part series. At the GTC, graphics performance was not the focus of Tesla Fermi. Rather the conference was emphasizing NVIDIA\u2019s new architecture as a revolutionary <em>G<\/em>eneral <em>P<\/em>urpose <em>P<\/em>rocessor that takes much more advantage of their new Fermi GPU\u2019s abilities of superfast parallel processing over their current architecture. NVIDIA\u2019s goal is to dominate the professional market with their Tesla GPUs.  Now that Fermi GF100 GPUs for NVIDIA\u2019s new video cards are finally in mass production, we will be looking at how NVIDIA intends to dominate gaming.<\/p>\n<div style=\"width: 311px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2009\/10\/fermprodmockup_thumb-jpg.jpg?resize=301%2C130\" alt=\"Fermi Production Mock-up\" data-recalc-dims=\"1\" \/><p class=\"wp-caption-text\">Fermi Production Mock-up<\/p><\/div>\n<div style=\"width: 195px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" src=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2009\/10\/rawgpu_ob_thumb-jpg.jpg?resize=185%2C128\" alt=\"Fermi GPU\" data-recalc-dims=\"1\" \/><p class=\"wp-caption-text\">Fermi GPU<\/p><\/div>\n<p>To summarize the new architecture, Fermi boasts a brand new shader core whose compute clusters comprise a single shader multiprocessor (SM). Each stream processor has a fully-pipelined integer arithmetic logic unit (ALU) and floating point unit (FPU). Each SM can dual-issue two independent instructions per clock to two different warps. Each instruction is run by a 16-way SIMD block that handles single-precision Floating Multiply-Add Instruction (FMAs). The Fermi memory hierarchy is also new, sporting a new unified L2 cache that serves all of the SMs without partitions. In addition, a new unified memory space allows each SM to not only communicate with its own local registers and shared memory, but now with L2 cache and beyond.<\/p>\n<p>The GF100 features 768KB unified level-two cache as well as a rather complex cache hierarchy. In addition, many other GPU-compute areas of performance are improved over NVIDIA\u2019s current Tesla architecture GPUs, GT200.  The GF100 hardware can sustain peak Single Precision (SP) and Double Precision (DP) FMA instruction throughput. Atomic instruction throughput is maximized over the current generation and Fermi is backed by ECC which is absolutely necessary for GPU computing. This all comes together to support a new type of multi-threading technology which improves the efficiency of the 512 cores working together. The entire Fermi family is compatible with DirectX 11, OpenGL 3.x and OpenCL 1.x application programming interfaces (APIs). The new chips are finally in mass production using 40nm process technology at TSMC.<\/p>\n<p>Let&#8217;s go ahead and see what is new and improved with GF100.<\/p>\n<p><!--nextpage--><!--pagetitle:GF100 Architecture--><strong><span style=\"color: #99ccff;\"><span style=\"text-decoration: underline;\">GF100 Architecture<\/span><\/span><\/strong><\/p>\n<p>Lets look at the diagrams:<\/p>\n<p><a href=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/architecture_1-jpg.jpg\"><img loading=\"lazy\" decoding=\"async\" style=\"border-bottom: 0px; border-left: 0px; display: inline; border-top: 0px; border-right: 0px\" title=\"Architecture_1\" src=\"http:\/\/i0.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/architecture_1_thumb-jpg.jpg?resize=244%2C199\" border=\"0\" alt=\"Architecture_1\" data-recalc-dims=\"1\" \/><\/a> <a href=\"http:\/\/i0.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/raster2-jpg.jpg\"><img loading=\"lazy\" decoding=\"async\" style=\"border-bottom: 0px; border-left: 0px; display: inline; border-top: 0px; border-right: 0px\" title=\"raster2\" src=\"http:\/\/i1.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/raster2_thumb-jpg.jpg?resize=244%2C198\" border=\"0\" alt=\"raster2\" data-recalc-dims=\"1\" \/><\/a> <a href=\"http:\/\/i1.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/dist_parallel-jpg.jpg\"><img loading=\"lazy\" decoding=\"async\" style=\"border-bottom: 0px; border-left: 0px; display: inline; border-top: 0px; border-right: 0px\" title=\"dist_parallel\" src=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/dist_parallel_thumb-jpg.jpg?resize=244%2C198\" border=\"0\" alt=\"dist_parallel\" data-recalc-dims=\"1\" \/><\/a><\/p>\n<p><em><span style=\"color: #99ccff;\">The first diagram from NVIDIA\u2019s slides, shows the GF100 block diagram illustrating the Host Interface, the GigaThread Engine, four GPCs, six Memory Controllers, six ROP partitions, and a 768 KB L2 cache. Each GPC contains four PolyMorph engines. The ROP partitions are immediately adjacent to the L2 cache. The second  image illustrates how GF100\u2019s graphics architecture is built from a number of hardware blocks called Graphics Processing Clusters (GPCs). A GPC contains a Raster Engine and up to four SMs. The third image illustrates how it all works together.<\/span><\/em><\/p>\n<p>Firstly, CPU commands are read by the GPU via the Host Interface. In turn the GigaThread Engine fetches data from the system memory and copies it to the framebuffer. GF100 implements six 64-bit GDDR5 memory controllers for 384-bit total which facilitates high bandwidth access to the framebuffer. The GigaThread Engine then creates and dispatches thread blocks to various SMs. Individual SMs in turn schedules warps (groups of 32 threads) to CUDA cores and to the other execution units. The GigaThread Engine also redistributes work to the SMs when work expansion occurs in the graphics pipeline.<\/p>\n<p>In the first image, the rectangular structures are SMs, or as NVIDIA calls them, streaming multiprocessors of which Fermi has sixteen. NVIDIA calls the green squares inside of each SM, &#8220;CUDA cores&#8221;. These CUDA cores compromise the chip&#8217;s most fundamental execution resource which helps to determine the chip&#8217;s total processing power and ultimately its performance. The GT200 has 240 and Fermi has 512.<\/p>\n<p>The memory interfaces are 64-bit.  This means that Fermi has its total path to memory that is 384 bits wide. This is in contrast to the higher 512 bit pathway on the GT200. However, Fermi compensates by delivering almost twice the bandwidth per pin due to its support for GDDR5 memory; GT200 used GDDR3 memory.<\/p>\n<p>To summarize, Fermi GF100 has:<\/p>\n<ul>\n<li><span style=\"color: #99ccff;\">512 CUDA cores<\/span><\/li>\n<li><span style=\"color: #99ccff;\">16 Geometry Units<\/span><\/li>\n<li><span style=\"color: #99ccff;\">4 raster units<\/span><\/li>\n<li><span style=\"color: #99ccff;\">64 texture units<\/span><\/li>\n<li><span style=\"color: #99ccff;\">48 ROP units<\/span><\/li>\n<li><span style=\"color: #99ccff;\">384-bit GDDR5<\/span><\/li>\n<\/ul>\n<p>NVIDIA\u2019s current generation product, the GT200 \u2013 of which GTX 285 is the single GPU flagship &#8211; was able to improve on the original G80 design as represented by the 8800 GTX. By refining G80\u2019s architecture, NVIDIA made it more programmable by adding double precision (DP) support and atomic operations.  GT200 managed all of this while still holding on to the highest performance crown for a single GPU until nearly five months ago when AMD\/ATI\u2019s Radeon 5870 launched. Their competitor has the first DX11 chip that was built with incremental changes made over its last generation resulting in significant performance improvements over HD 4800 series.<\/p>\n<p>So now NVIDIA has announced their Fermi GF100 next generation DX11 architecture which aims for even greater performance and also is more programmable and software friendly. There is no \u201cGT300\u201d. Until now, NVIDIA has chosen to primarily discuss Fermi Tesla GPU computing architecture and not to disclose microarchitecture or especially game-related performance details of GF100.<\/p>\n<p>The biggest changes in GF100 architecture show us that the geometry pipeline has been significantly revamped with improved performance in geometry shading, stream out, and culling. Fillrate has also been improved which enables multiple displays to be driven simultaneously by GF100 SLI, much like AMD\u2019s Eyefinity; but now additionally in 3D and at 120 Hz.<\/p>\n<p>From studying the second image, we can see that the GPC is GF100\u2019s dominant high-level hardware block. It features two key innovations\u2014a scalable Raster Engine for triangle setup, rasterization, and z-cull, and a scalable PolyMorph Engine for vertex attribute fetch and tessellation. The Raster Engine resides in the GPC, whereas the PolyMorph Engine resides in the SM. On earlier NVIDIA GPUs, SMs and Texture Units were grouped together in hardware blocks called Texture Processing Clusters (TPCs). On GF100, each SM has four dedicated Texture Units.<\/p>\n<p>As we look deeper, we can see that Fermi&#8217;s tessellation engine is impressive. It is not something just \u201ctacked on\u201d to GT200. NVIDIA saw early on that if they only made incremental changes to GT200, they would run into severe bottlenecks. Simply adding tessellation to GT200 would lead to intolerable geometry bottlenecks. They tell us that this is what took them so long \u2013 they had to design a better balanced new chip architecture that could also have better sequential rendering semantics built into its engine.<\/p>\n<p><!--nextpage--><!--pagetitle:Geometry--><strong><span style=\"color: #99ccff;\"><span style=\"text-decoration: underline;\">Geometry<\/span><\/span><\/strong><\/p>\n<p>NVIDIA\u2019s goal is for GF100 to enable film-like geometric realism for game characters and objects. Geometric realism is central to the GF100 architectural enhancements for graphics. In addition, PhysX simulations are faster and developers can utilize GPU computing features in games more easily and effectively.<\/p>\n<p>While programmable shading has allowed PC games to mimic the cinema in per-pixel effects, geometric realism is way behind. The most advanced modern PC games will use one to two million polygons per frame whereas a typical frame in a computer generated film uses hundreds of millions of polygons. While the number of pixel shaders has grown from one to many hundreds, the triangle setup engine has remained a singular unit. For example, the GeForce GTX 285 has more than 150 times the shading horsepower of the old GeForce FX, but less than 3 times the geometry processing rate. This means that pixels are shaded well but their geometric detail is weak.<\/p>\n<p>Take a look at NVIDIA&#8217;s example from <em>Far Cry 2<\/em>. The holster has a heavily segmented strap. The corrugated roof is just a flat surface with a striped texture instead of curving properly. We also note that this character wears a hat to avoid the complexity of rendering hair.<\/p>\n<p style=\"TEXT-ALIGN: center\"><a href=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/gameslackgeometry.PNG\"><img loading=\"lazy\" decoding=\"async\" class=\"size-medium wp-image-18032 aligncenter\" title=\"GamesLackGeometry\" src=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/gameslackgeometry.PNG?resize=300%2C112\" alt=\"GamesLackGeometry\" srcset=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/gameslackgeometry.PNG?resize=300%2C112 300w, http:\/\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/gameslackgeometry.PNG 640w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" data-recalc-dims=\"1\" \/><\/a><\/p>\n<p><span style=\"color: #99ccff;\"><em>On the other hand, the exquisitely detailed characters in CG films are made possible by tessellation and displacement mapping. Tessellation refines large triangles into collections of smaller triangles, while displacement mapping changes their relative position. To achieve these same goals, GF100\u2019s entire graphics pipeline is designed to deliver higher performance in tessellation and geometry throughput.<\/em><\/span><\/p>\n<p>GF100 replaces the traditional geometry processing architecture at the front end of the graphics pipeline with an entirely new distributed geometry processing architecture that is implemented using multiple \u201cPolyMorph Engines\u201d. Each of these engine includes a tessellation unit, an attribute setup unit, and other geometry processing units. Each SM has its own dedicated PolyMorph Engine as shown by the three grouped diagrams that we showed you earlier (above).<\/p>\n<p>Newly generated primitives are converted to pixels by four Raster Engines that operate in parallel compared to a single Raster Engine in GT200 and in earlier GPUs. On-chip L1 and L2 caches now enable high bandwidth transfer of primitive attributes between the SM and the tessellation unit as well as between different SMs. Tessellation and all its supporting stages are performed in parallel on GF100 with improved geometry throughput. GF100\u2018s ability to perform parallel geometry processing is possibly the single most important GF100 architectural improvement. The ability to deliver setup rates exceeding one primitive per clock while maintaining correct rendering order is a significant technical achievement.<\/p>\n<p>Major compute features improved on GF100 that will be useful in games include faster context switching between graphics and PhysX, concurrent compute kernel execution and an enhanced caching architecture which is good for irregular algorithms such as ray tracing, and AI. Simultaneously, improved atomic operations performance allows threads to safely cooperate through work queues, accelerating novel rendering algorithms. For example, fast atomic operations allow transparent objects to be rendered without presorting (order independent transparency) enabling developers to create levels with complex glass environments. GF100\u2019s GigaThread engine reduces context switch time, making it possible to execute multiple compute and physics kernels for each frame.<\/p>\n<p><!--nextpage--><!--pagetitle:Tessellation &#038; Displacement Mapping--><strong><span style=\"color: #99ccff;\"><span style=\"text-decoration: underline;\">Tessellation and Displacement Mapping<\/span><\/span><\/strong><\/p>\n<p>It takes DX11 to take advantage of geometry. DX9 and DX10 are unable to create generalized geometry on the GPU. Therefore we will see Tessellation and displacement mapping used together to create more realism in games. The ability to control the geometric level of detail (LOD) is very important. Because it is on-demand and the data is all kept on-chip, precious memory bandwidth is preserved. Also, because one model may produce many LODs, the same game assets may be used on a variety of platforms which makes the game developers very happy. Their characters can also be easily adjusted as to how it appears in the scene; if it is small then it gets little geometry, if it is close to the screen then it is rendered with greater detail.<\/p>\n<p>As an additional benefit, developers may be able to use the same models on many generations of games and future GPUs where performance increases will allow for enabling even greater detail than was possible when the game was first released. Complexity can be adjusted dynamically to even target a given frame rate!<\/p>\n<p>Here in NVIDIA\u2019s slide from the Unigine engine demo, we see tessellation compared, on and off. There is no comparison; tessellation adds to realism.<\/p>\n<p style=\"TEXT-ALIGN: center\"><a href=\"http:\/\/i0.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/tesselation-jpg.jpg\"><img loading=\"lazy\" decoding=\"async\" style=\"border-bottom: 0px; border-left: 0px; display: inline; border-top: 0px; border-right: 0px\" title=\"Tesselation\" src=\"http:\/\/i0.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/tesselation_thumb-jpg.jpg?resize=244%2C76\" border=\"0\" alt=\"Tesselation\" data-recalc-dims=\"1\" \/><\/a><\/p>\n<p>Take a look at the third image that we presented earlier in this article. The use of tessellation fundamentally changes the GPU\u2019s graphics workload balance. With tessellation, the triangle density of a given frame can increase by multiple orders of magnitude which strains serial resources such as the setup and rasterization units. To facilitate high triangle rates, NVIDIA designed a scalable geometry engine called the PolyMorph Engine. Each of GF100\u2019s 16 PolyMorph engines has its own dedicated vertex fetch unit and a tessellator which expands geometry performance.<\/p>\n<p>In conjunction with the PolyMorph Engine, NVIDIA designed four parallel Raster Engines which allows up to four triangles to be setup per clock. Results calculated in each of five stages which are then passed to an SM. The SM executes the game\u2019s shader, returning the results to the next stage in the PolyMorph Engine. After all stages are complete, the results are forwarded to one of the four Raster Engines.<\/p>\n<p>The Rasterizer takes the edge equations for each primitive and computes pixel coverage. If antialiasing is enabled, coverage is performed for each multisample and coverage sample. Each Rasterizer outputs eight pixels per clock for a total of 32 rasterized pixels per clock across the chip. Pixels produced by the rasterizer are sent to the Z-cull unit. By having a dedicated tessellator for each SM, and a Raster Engine for each GPC, GF100 delivers up to 8 times the geometry performance of GT200. NVIDIA also compares the geometry performance of GF100 to HD 5870 and finds Fermi is significantly faster.<\/p>\n<p>Here is a performance comparison between GF100 and HD 5870 using a 60 second run with the Unigine engine:<\/p>\n<p style=\"TEXT-ALIGN: center\"><a href=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/60secunigine-jpg.jpg\"><img loading=\"lazy\" decoding=\"async\" style=\"border-bottom: 0px; border-left: 0px; display: inline; border-top: 0px; border-right: 0px\" title=\"60SecUnigine\" src=\"http:\/\/i1.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/60secunigine_thumb-jpg.jpg?resize=244%2C180\" border=\"0\" alt=\"60SecUnigine\" data-recalc-dims=\"1\" \/><\/a><\/p>\n<p><!--nextpage--><!--pagetitle:Anti-Aliasing--><strong><span style=\"color: #99ccff;\"><span style=\"text-decoration: underline;\">Anti-Aliasing Image Quality<\/span><\/span><\/strong><\/p>\n<p>To improve anti-aliasing image quality, the GF100 introduces a new anti-aliasing mode: 32xCSAA. nVidia\u2019s previous strongest edge AA mode was 16xQ, but this is now bested by 32xAA. Here\u2019s the sample pattern for it, courtesy of nVidia:<\/p>\n<p style=\"TEXT-ALIGN: center\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-14572\" title=\"32\" src=\"http:\/\/i1.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/32.png?resize=185%2C181\" alt=\"32\" data-recalc-dims=\"1\" \/><\/p>\n<p><span style=\"color: #99ccff;\">32x<\/span>AA = <span style=\"color: #99ccff;\">8x<\/span>MSAA + <span style=\"color: #99ccff;\">24x<\/span>CSAA.<\/p>\n<p>Thus 32xCSAA is a natural extension of 16xQ, and offers even stronger edge (polygon) anti-aliasing, courtesy of providing a total of 32 unique samples<\/p>\n<p>But that\u2019s not all that has improved. The GF100 has a new ability to use coverage samples to affect the quality of alpha textures, as implemented through transparency anti-aliasing. With previous nVidia hardware such as the GT200, coverage samples had no effect on transparency anti-aliasing quality, as the result was derived solely from the base multi-sampling pattern in effect.<\/p>\n<p>Also in the specific case of transparency multi-sampling, image quality has improved there too. Any titles using the older alpha test method to render transparent textures have their shader code automatically converted to use the alpha-to-cover technique, which should greatly improve image quality, especially in heavily aliased areas.<\/p>\n<p>The upshot of this is higher quality edges, and higher quality alpha textures.<\/p>\n<p><strong> <\/strong><\/p>\n<p><strong><span style=\"color: #99ccff;\"><span style=\"text-decoration: underline;\">Anti-Aliasing Performance<\/span><\/span><\/strong><\/p>\n<p>In addition to improving image quality, anti-aliasing performance has also increased. When it comes to AA, the most obvious area to target is the ROPs, and that\u2019s exactly what nVidia has done. The GF100 has <span style=\"color: #99ccff;\">48<\/span> ROPs, up from <span style=\"color: #99ccff;\">32<\/span> ROPs on the GTX285, which is especially helpful for portions of the scene that cannot be compressed.<\/p>\n<p>Each ROP is also faster and more efficient than on previous generations, so it can do more work per cycle. This includes improvements made to the compression technology.<\/p>\n<p style=\"TEXT-ALIGN: center\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-14568\" title=\"25a\" src=\"http:\/\/i1.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/25a.png?resize=357%2C357\" alt=\"25a\" srcset=\"http:\/\/i1.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/25a.png?resize=357%2C357 357w, http:\/\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/25a-150x150.png 150w, http:\/\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/25a-300x300.png 300w\" sizes=\"auto, (max-width: 357px) 100vw, 357px\" data-recalc-dims=\"1\" \/><\/p>\n<p>Aside from better AA performance in general, nVidia\u2019s old Achilles heel with 8xMSAA performance should also be addressed by the improvements. Historically, prior nVidia architectures have exhibited much higher relative performance hit when going from 4xMSAA to 8xMSAA, compared to competing ATi architectures.<\/p>\n<p>Also by moving to 384 bit GDDR5, nVidia should have access to plenty of memory bandwidth to keep all of those ROPs fed with data.<\/p>\n<p><!--nextpage--><!--pagetitle:Texture Filtering--><span style=\"color: #99ccff;\"><strong><span style=\"text-decoration: underline;\">Texture Filtering<\/span><\/strong><\/span><\/p>\n<p>As with anti-aliasing, there have been improvements made to texturing too. Interestingly the GF100 only has <span style=\"color: #99ccff;\">64<\/span> TMUs, which is much less than the <span style=\"color: #99ccff;\">80<\/span> TMUs on the GTX285, but nVidia claims overall performance should still be higher because of improvements to performance and efficiency.<\/p>\n<p>Texture caching has been substantially improved, with the L1 cache being redesigned for greater efficiency. Also the presence of a unified L2 cache means the texture cache size is three times higher than on the GT200.<\/p>\n<p>Layout changes and internal improvements to the texture units also combine with a higher TMU clock speed. On the GT200 the TMUs ran at the GPU\u2019s core clock; on the GF100 they run at a higher clock, which allows them to perform more work in the same amount of time. nVidia\u2019s numbers show 40% to 70% higher texturing performance than the GT200, despite having much fewer TMUs.<\/p>\n<p style=\"TEXT-ALIGN: center\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-14569\" title=\"25b\" src=\"http:\/\/i1.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/25b.png?resize=404%2C369\" alt=\"25b\" srcset=\"http:\/\/i1.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/25b.png?resize=404%2C369 404w, http:\/\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/25b-300x274.png 300w\" sizes=\"auto, (max-width: 404px) 100vw, 404px\" data-recalc-dims=\"1\" \/><\/p>\n<p>The GF100\u2019s texture units also offer hardware accelerated jittered sampling. This essentially means the hardware has the ability to offer a form of stochastic filtering by varying the texture sampling on a per-pixel basis. This is done by implementing DirectX 11\u2019s Gather4 in hardware, and it provides the ability for up to four texels to be fetched from a 128&#215;128 pixel grid with a single instruction.<\/p>\n<p style=\"TEXT-ALIGN: center\"><a href=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/29.png\" target=\"_blank\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-thumbnail wp-image-14571\" title=\"29\" src=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/29.png?resize=150%2C150\" alt=\"29\" data-recalc-dims=\"1\" \/><\/a><\/p>\n<p>This not only improves performance with things like ambient occlusion, but it can also improve image quality by removing banding through random sampling. It also allows game developers to implement customized texture filtering more efficiently. nVidia states that the GF100\u2019s hardware implementation of this technique offers up to twice the performance of the GT200.<\/p>\n<p><!--nextpage--><!--pagetitle:Compute Architecture &#038; Ray Tracing--><span style=\"color: #99ccff;\"><strong><span style=\"text-decoration: underline;\">Compute Architecture<\/span><\/strong><\/span><\/p>\n<p>The compute engine is designed to handle the GPGPU side of things and encapsulates features such as CUDA, OpenCL, Direct Compute, and PhysX. Many of these have been around since the G80 days, but the GF100 delivers a number of improvements to make such general purpose computing run better.<\/p>\n<p>The GF100 is designed to handle a wider range of algorithms better to encourage the use of the GPU more for parallel problems. One key area of improvement comes from its better cache system, which allows threads that access the same memory locations to run faster.<\/p>\n<p>Another key improvement allows the GF100 to execute multiple task kernels at once, and the context switching between such tasks is much faster than on previous GPUs. This differs from the GT200 which could only run one task kernel at a time, and had very slow context switching.<\/p>\n<p>And lastly, high level features such as debugging and a C++ programming environment to access GPGPU features are made possible with nVidia\u2019s <span style=\"color: #99ccff;\">Nexus<\/span> plug-in for Visual Studio. Such features simplify programming GPGPU tasks as they assist developers to work at a higher level than was previously possible.<\/p>\n<p><strong> <\/strong><\/p>\n<p><span style=\"color: #99ccff;\"><strong><span style=\"text-decoration: underline;\">Ray Tracing<\/span><\/strong><\/span><\/p>\n<p style=\"TEXT-ALIGN: center\"><a href=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/rt-jpg.jpg\"><img loading=\"lazy\" decoding=\"async\" style=\"border-bottom: 0px; border-left: 0px; display: inline; border-top: 0px; border-right: 0px\" title=\"RT\" src=\"http:\/\/i0.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/rt_thumb-jpg.jpg?resize=244%2C154\" border=\"0\" alt=\"RT\" data-recalc-dims=\"1\" \/><\/a><\/p>\n<p>The GF100 will not be able to do complex ray tracing (RT) in real time in PC games as in the above image. However, NVIDIA believes that RT is the future of graphics and they expect some implementation of it in conjunction with rasterization fairly soon as developers begin to take advantage of GF100\u2019s new programming capabilities.<\/p>\n<p><!--nextpage--><!--pagetitle:Conclusion--><span style=\"color: #99ccff;\"><strong><span style=\"text-decoration: underline;\">Conclusion<\/span><\/strong><\/span><\/p>\n<p>It\u2019s clear that nVidia has invested a lot of resources and design effort into trying to make the GF100 the fastest single GPU to date. In addition to a very clear focus on improving GPGPU performance and usability, numerous enhancements to image quality and performance for gaming purposes have also been made.<\/p>\n<p>It\u2019ll be very interesting to see how the card performs in actual gaming situations, and more importantly, how it compares to ATi\u2019s current single GPU flagship, the Radeon 5870.<\/p>\n<p>We are looking forward to bringing our readers the latest news about the Fermi GF100 and we will be testing its performance and image quality in gaming. There is much more to be revealed about NVIDIA&#8217;s new GPU. Stay tuned.  The graphics wars are heating up and it is getting very interesting again.<\/p>\n<p style=\"TEXT-ALIGN: center\"><a href=\"http:\/\/i2.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/turbulence-jpg.jpg\"><img loading=\"lazy\" decoding=\"async\" style=\"border-bottom: 0px; border-left: 0px; display: inline; border-top: 0px; border-right: 0px\" title=\"Turbulence\" src=\"http:\/\/i0.wp.com\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/turbulence_thumb-jpg.jpg?resize=244%2C196\" border=\"0\" alt=\"Turbulence\" data-recalc-dims=\"1\" \/><\/a><\/p>\n<p><strong> <\/strong><\/p>\n<p>Article written by <span style=\"color: #00ff00;\"><strong>Mark Poppin<\/strong><\/span> and <span style=\"color: #00ff00;\"><strong>BFG10K<\/strong><\/span>, AlienBabelTech Senior Editors.<\/p>\n<p><strong> <\/strong><\/p>\n<blockquote><p><span style=\"color: #99ccff;\">Please join us in our <a href=\"http:\/\/alienbabeltech.com\/abt\/index.php\" target=\"_blank\">Forums<\/a><\/span><\/p>\n<p><span style=\"color: #99ccff;\">Follow us on <a href=\"http:\/\/twitter.com\/alienbabeltech\" onclick=\"_gaq.push(['_trackEvent', 'outbound-article', 'http:\/\/twitter.com\/alienbabeltech', 'Twitter']);\" target=\"_blank\">Twitter<\/a><\/span><\/p>\n<p><span style=\"color: #99ccff;\">For the latest updates from ABT, please <a href=\"http:\/\/alienbabeltech.com\/main\/?feed=rss2\" target=\"_blank\">join our RSS News Feed<\/a><\/span><\/p><\/blockquote>\n","protected":false},"excerpt":{"rendered":"<p><a href=\"http:\/\/alienbabeltech.com\/main\/?p=14600\"><img decoding=\"async\" align=top title=\"Author: BFG10K &#038; Apoppin\" src=\"http:\/\/alienbabeltech.com\/main\/wp-content\/uploads\/2010\/01\/Article-Image.png\"><\/a><\/p>\n","protected":false},"author":6,"featured_media":18032,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2533,7814,7],"tags":[1471,2354,2362,811,121],"class_list":["post-14600","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-abt-news","category-articles","category-technology","tag-directx-11","tag-fermi","tag-gf100","tag-gt300","tag-nvidia"],"_links":{"self":[{"href":"http:\/\/alienbabeltech.com\/main\/wp-json\/wp\/v2\/posts\/14600","targetHints":{"allow":["GET"]}}],"collection":[{"href":"http:\/\/alienbabeltech.com\/main\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"http:\/\/alienbabeltech.com\/main\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"http:\/\/alienbabeltech.com\/main\/wp-json\/wp\/v2\/users\/6"}],"replies":[{"embeddable":true,"href":"http:\/\/alienbabeltech.com\/main\/wp-json\/wp\/v2\/comments?post=14600"}],"version-history":[{"count":0,"href":"http:\/\/alienbabeltech.com\/main\/wp-json\/wp\/v2\/posts\/14600\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"http:\/\/alienbabeltech.com\/main\/wp-json\/wp\/v2\/media\/18032"}],"wp:attachment":[{"href":"http:\/\/alienbabeltech.com\/main\/wp-json\/wp\/v2\/media?parent=14600"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"http:\/\/alienbabeltech.com\/main\/wp-json\/wp\/v2\/categories?post=14600"},{"taxonomy":"post_tag","embeddable":true,"href":"http:\/\/alienbabeltech.com\/main\/wp-json\/wp\/v2\/tags?post=14600"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}