Posts Tagged ‘vmmark’

h1

Quick Take: HP’s Sets Another 48-core VMmark Milestone

August 26, 2009

Not satisfied with a landmark VMmark score that crossed the 30 tile mark for the first time, HP’s performance team went back to the benches two weeks later and took another swing at the performance crown. Well, the effort paid off, and HP significantly out-paced their two-week-old record with a score of 53.73@35 tiles in the heavy weight, 48-core category.

Using the same 8-processor HP ProLiant DL785 G6 platform as in the previous run – complete with 2.8GHz AMD Opteron 8439 SE 6-core chips and 256GB DDR2/667 – the new score comes with significant performance bumps in the javaserver, mailserver and database results achieved by the same system configuration as the previous attempt – including the same ESX 4.0 version (164009). So what changed to add an additional 5 tiles to the team’s run? It would appear that someone was unsatisfied with the storage configuration on the mailserver run.

Given that the tile ratio of the previous run ran about 6% higher than its 24-core counterpart, there may have been a small indication that untapped capacity was available. According to the run notes, the only reported changes to the test configuration – aside from the addition of the 5 LUNs and 5 clients needed to support the 5 additional tiles – was a notation indicating that the “data drive and backup drive for all mailserver VMs” we repartitioned using AutoPart v1.6.

The change in performance numbers effectively reduces the virtualization cost of the system by 15% to about $257/VM – closing-in on its 24-core sibling to within $10/VM and stretching-out its lead over “Dunnington” rivals to about $85/VM. While virtualization is not the primary application for 8P systems, this demonstrates that 48-core virtualization is definitely viable.

SOLORI’s Take: HP’s performance team has done a great job tuning its flagship AMD platform, demonstrating that platform performance is not just related to hertz or core-count but requires balanced tuning and performance all around. This improvement in system tuning demonstrates an 18% increase in incremental scalability – approaching within 3% of the 12-core to 24-core scaling factor, making it actually a viable consideration in the virtualization use case.

In recent discussions with AMD about the SR5690 chipset applications for Socket-F, AMD re-iterated that the mainstream focus for SR5690 has been Magny-Cours and the Q1/2010 launch. Given the close relationship between Istanbul and Magny-Cours – detailed nicely by Charlie Demerjian at Semi-Accurate – the bar is clearly fixed for 2P and 4P virtualization systems designed around these chips. Extrapolating from the similarities and improvements to I/O and memory bandwidth, we expect to  see 2P VMmarks besting 32@23 and 4P scores over 54@39 from HP, AMD and Magny-Cours.

SOLORI’s 2nd Take: Intel has been plugging away with its Nehalem-EX for 8-way systems and – delivering 128-threads – promises to deliver some insane VMmarks. Assuming Intel’s EX scales as efficiently as AMD’s new Opterons have, extrapolations indicate performance for the 4P, 64-thread Nehalem-EX shoud fall between 41@29 and 44@31 given the current crop of speed and performance bins. Using the same methods, our calculus predicts an 8P, 128-thread EX system should deliver scores between 64@45 and 74@52.

With EX expected to clock at 2.66GHz with 140W TDP and AMD’s MCM-based Magny-Cours doing well to hit 130W ACP in the same speed bins, CIO’s balancing power and performance considerations will need to break-out the spreadsheets to determine the winners here. With both systems running 4-channel DDR3, there will be no power or price advantage given on either side to memory differences: relative price-performance and power consumption of the CPU’s will be major factors. Assuming our extrapolations are correct, we’re looking at a slight edge to AMD in performance-per-watt in the 2P segment, and a significant advantage in the 4P segment.

h1

Quick Take: HP Plants the Flag with 48-core VMmark Milestones

August 12, 2009

Following on the heels of last month we predicted that HP could easily claim the VMmark summit with its DL785 G6 using AMD’s Istanbul processors:

If AMD’s Istanbul scales to 8-socket at least as efficiently as Dunnington, we should be seeing some 48-core results in the 43.8@30 tile range in the next month or so from HP’s 785 G6 with 8-AMD 8439 SE processors. You might ask: what virtualization applications scale to 48-cores when $/VM is doubled at the same time? We don’t have that answer, and judging by Intel and AMD’s scale-by-hub designs coming in 2010, that market will need to be created at the OEM level.

Well, HP didn’t make us wait too long. Today, the PC maker cleared two significant VMmark milestones: crossing the 30 tile barrier in a single system (180 VMs) and exceeding the 40 mark on VMmark score. With a score of 47.77@30 tiles, the HP DL785 G6 – powered by 8 AMD Istanbul 8439 SE processors and 256GB of DDR2/667 memory – set the bar well beyond the competition and does so with better performance than we expected – most likely due to AMD’s “HT assist” technology increasing its scalability.

Not available until September 14, 2009, the HP DL785 G6 is a pricey competitor. We estimate – based on today’s processor and memory prices – that a system as well appointed as the VMmark-configured version (additional NICs, HBA, etc) will run at least $54,000 or around $300/VM (about $60/VM higher than the 24-core contender and about $35/VM lower than HP’s Dunnnigton “equivalent”).

SOLORI’s Take: While the September timing of the release might imply a G6 with AMD’s SR5690 and IOMMU, we’re doubtful that the timing is anything but a coincidence: even though such a pairing would enable PCIe 2.0 and highly effective 10Gbps solutions. The modular design of the DL785 series – with its ability to scale from 4P to 8P in the same system – mitigates the economic realities of the dwindling 8P segment, and HP has delivered the pinnacle of performance for this technology.

We are also impressed with HP’s performance team and their ability to scale Shanghai to Istanbul with relative efficiency. Moving from DL785 G5 quad-core to DL785 G6 six-core was an almost perfect linear increase in capacity (95% of theoretical increase from 32-core to 48-core) while performance-per-tile increased by 6%. This further demonstrates the “home run” AMD has hit with Istanbul and underscores the excellent value proposition of Socket-F systems over the last several years.

Unfortunately, while they demonstrate a 91% scaling efficiency from 12-core to 24-core, HP and Istanbul have only achieved a 75% incremental scaling efficiency from 24-cores to 48-cores. When looking at tile-per-core scaling using the 8-core, 2P system as a baseline (1:1 tile-to-core ratio), 2P, 4P and 8P Istanbul deliver 91%, 83% and 62.5% efficiencies overall, respectively. However, compared to the %58 and 50% tile-to-core efficiencies of Dunnington 4P and 8P, respectively, Istanbul clearly dominates the 4P and 8P performance and price-performance landscape in 2009.

In today’s age of virtualization-driven scale-out, SOLORI’s calculus indicates that multi-socket solutions that deliver a tile-to-core ratio of less than 75% will not succeed (economically) in the virtualization use case in 2010, regardless of socket count. That said – even at a 2:3 tile-to-core ratio – the 8P, 48-core Istanbul will likely reign supreme as the VMmark heavy-weight champion of 2009.

SOLORI’s 2nd Take: HP and AMD’s achievements with this Istanbul system should be recognized before we usher-in the next wave of technology like Magny-Cours and Socket G34. While the DL785 G6 is not a game changer, its footnote in computing history may well be as a preview of what we can expect to see out of Magny-Cours in 2H/2010. If 12-core, 4P system price shrinks with the socket count we could be looking at a $150/VM price-point for a 4P system: now that would be a serious game changer.

h1

NEC Adds Top 48-Core, Dell Challenges 24-Core in VMmark Race

July 29, 2009

NEC’s venerable Express5800/A1160 tops the 48-core VMmark category today with a score of 34.05@24 tiles to wrest the title away from IBM who established the category back in June, 2009. NEC’s new “Dunnington” X7460 Xeon-based score represents a performance per tile ratio of 1.41 and a tile to core efficiency of 50% using 128GB of ECC DDR2 RAM.

Compared to the leading 24-core “Dunnington” results – held by IBM’s x3850 M2 at 20.41@14 tiles – the NEC benchmark sets a scalability factor of 85.7% when moving from 4-socket to 8-socket systems. Both servers from NEC and IBM are scalable systems allowing for multiple chassis to be interconnected to achieve greater CPU-per-system numbers – each scaling in 4-CPU increments – ostensibly for OLTP advantages. The NEC starts at around $70K for 128GB and 48-cores resulting in a $486/VM cost to VMmark.

Also released today, Dell’s PowerEdge R905 – with 24 2.8GHz Istanbul cores (8439 SE) and 128GB of ECC DDR2 RAM – secures the number two slot in the 24-category with a posting of 29.51@20 tiles. This represents a tile ratio of 1.475 and tile efficiency of 83.3% for the $29K rack server from Dell at about $240/VM. Compared to its 12-core counterpart, this represents a 91% scalability factor.

If AMD’s Istanbul scales to 8-socket at least as efficiently as Dunnington, we should be seeing some 48-core results in the 43.8@30 tile range in the next month or so from HP’s 785 G6 with 8-AMD 8439 SE processors. You might ask: what virtualization applications scale to 48-cores when $/VM is doubled at the same time? We don’t have that answer, and judging by Intel and AMD’s scale-by-hub designs coming in 2010, that market will need to be created at the OEM level.

Based on the performance we’re seeing in 8-socket systems relative to 4-socket and the upcoming “massively mult-core” processors in 2010, the law of diminishing returns seems to favor the 4-socket system as the limit for anything but massive OLTP workloads. Even then, we expect to see 48-core in a “4-way” box more efficient than the same number of cores in an 8-way box. The choice in virtualization will continue to be workload biased, with 2P systems offering the best “small footprint” $/VM solution and 4P systems offering the best “large footprint” $/VM solution.

h1

RIP Dunnington: HP’s 4P/24-core Istanbul Takes VMmark Summit

July 15, 2009
HP has simultaneously achieved two near identical VMmark scores with their ProLiant DL585 G6 rack server and ProLiant BL685c G6 blade, claiming the summit from the reigning 24-core champion. Since first establishing the 24-core tier VMmark in September 2009, the Intel “Dunnington” 6-core processor (FSB architecture) has gone unchallenged. Now, with the release of the Opteron 8439SE raising the performance bar and the Opteron 8435 making a clear price-performance case, Dunnington’s vacation is over.

Today’s Istanbul-based achievements – established in the same memory footprint as the top Dunnington – renders the venerable processor all but obsolete, besting the champ by 4 tiles (24 more virtual machines) with a score-tile ratio of 1.5 for the rack system and 1.46 (same as the Dunnington at 14 tiles) for the blade. Using the HP and IBM on-line configuration tools, we established the retail (on-line) price for each system – down to the Fiber Channel HBA’s – and compared them for $/VM value. Here are the results:

HP DL685 G6 HP BL685c G6 IBM x3850 M2
Processor 4x Opteron 8439SE 2.8GHz 4x Opteron 8435 2.6GHz 4x Xeon X7460 2.67GHz
Memory 128GB (16x8GB PC2-5300 Reg ECC) 128GB (16x8GB PC2-5300 Reg ECC) 128GB (32x4GB PC2-5300 Reg ECC)
LAN Controllers 1x Dual-Port NC371i 1Gbps,
3x Dual-Port NC380T 1Gbps
2x Dual-Port NC532i Flex-10 10Gbs,
1x Dual-Port NC360m 1Gbps
2x Intel PRO 1000PT Dual-Port 1Gbps
HBA Qlogic QMH2462 Dual-Port FC Qlogic QMH2462 Dual-Port FC 2x Qlogic QMH2462 Dual-Port FC
OS RAID Controller HP Smart Array P800 HP Smart Array P400i HBA
OS Disks 2x 73Gb SAS 10K 2x 73Gb SAS 10K SAN
On-line Price $36,862.00 $35,296.00 $34,269.00
On-line w/3rd Party Memory $28,712.00 $27,356.00 $33,207.00
VMmark Results 29.95@20 tiles 29.19@20 tiles 20.5@14 tiles
VMmark Tile Ratio 1.5 1.46 1.46
Cost/VM Retail $307.18 $294.13 $407.96
Cost/VM 3rd Party $239.27 $227.97 $276.73

The results indicate a 21-38% savings per-VM for Istanbul over Dunnington in the 4P/24-core virtualization space. This is bread-and-butter territory for VDI implementations and SQL virtualizations, and Intel’s last remaining market place for the Dunnington processor. With the top-bin Istanbul weighing-in with 3% better performance, 18% less power consumption and 30% more capacity against Dunnington at the same price point, Intel’s 4P gambit is played-out and Nehalem-EX cannot arrive too soon for Intel.

It is worth asking the question: does the HP ProLiant 4P/24-core offer the best value? The answer depends on the value proposition. From a straight $/VM vantage point, the HP DL385 G6 comparison demonstrated a more economical $182/VM – a difference of $40/VM lower than the BL685c G6 – so the 2P rack system still comes out on top for the absolute bottom-line concious. However, for applications like SQL consolidations, the additional savings in licensing on 4P platforms versus 2P platforms dwarfs this differential.

What is clear: AMD’s Istanbul solution will remain unchallenged in the 4P space both in raw performance and in price-performance until Nehalem-EX is delivered. That means if Nehalem-EX does not arrive in Q3/2009, the market will likely wait for Q1/2010 to make any long-term purchasing decisions in anticipation of the new platforms slated to break-in the new year.
h1

Lenovo Claims Top VMmark Spot: 2P, 8C

July 1, 2009

The new top spot for VMmark in the “8 core” category is now held by Lenovo’s R525 G2 rack server with a score of 24.35@17 tiles (tile ratio of 1.43 over 102 VMs). As this server appears to be available in the overseas (China) markets only, we can only estimate the street price of the system used in the benchmark based on the reported build-out at to be around $20,330 per server (street):

  • Base Lenovo R525 G2 ($4,900 – 30,000 yuan)
  • 2 x Intel Xeon X5570 Processors ($1,500/ea)
  • 96GB ECC DDR3/1066 (12x8GB) ($900/DIMM from Kingston)
  • 1 x Intel 82575EB dual-port GigabitEthernet (on-board)
  • 2 x Intel 82571EB dual-port GigabitEthernet (2x PCIe slot, $150/ea)
  • 1 x QLogic QLE2462 FC HBA (1x PCIe slot, $1,300)
  • 1 x LSI1078 SAS Controller (on-board)
  • 2 x SAS OS drive ($300 est.)

An EMC CX3-40f was used as the storage backing of the test. The storage system included 4GB cache, 4 enclosures and 55 146GB 15K FC disks (10, 15, 15, 15), and 17 LUNs at 100GB each. Interestingly, a Cisco Linksys SR2024 GigabitEthernet switch was used for the network interconnection (about $299/each at NewEgg) which implies that test results are not being influenced on network performance or latency. Given the use of a 2-port FC HBA for storage, iSCSI network performance is not a factor.

At about $1,094/tile ($182/VM) the new “top dog” delivers its best at a 5% price-per-VM premium over Istanbul’s only VMmark results (1.41 tile ratio) and an 80% system price premium (assuming memory sourced by third parties).  Since we had to go to the street to configure the Lenovo system, the Istanbul system saves about $1,570 [in mark-up] under similar (non-vendor pricing) circumstances:

  • Base HP DL385 G6  ($5,100)
  • 2 x AMD 2435 Istanbul Processors (included)
  • 64GB ECC DDR2/800 (8x8GB) ($370/DIMM)
  • 2 x Broadcom 5709 dual-port GigabitEthernet (on-board)
  • 1 x Intel 82571EB dual-port GigabitEthernet (1x PCIe slot, $150/ea)
  • 1 x QLogic QLE2462 FC HBA (1x PCIe slot, $1,300)
  • 1 x HP SAS Controller (on-board)
  • 2 x SAS OS drive (included)
  • $9,810/system total (versus $11,378 complete from HP)

Street pricing changes Istanbul’s numbers to $892/tile ($149/VM) signifying a 22% per-VM savings and a 52% savings in system price. Given that virtualization systems are generally sold in pairs, this comparison shows that a redundant Istanbul system can be had for less than the cost of a non-redundant Nehalem. For SMB’s getting started in virtualization, Istanbul continues to offer a compelling system value proposition over Nehalem.

h1

Dell Posts Top 4P/16-core VMmark

June 21, 2009

Dell has posted a new VMmark for its PowerEdge M905 series with a score of 22.90@17 tiles for a 4P Opteron 8393 SE based system. Although newly posted on VMware’s VMmark scoreboard, this test was performed on ESX 4.0 (build 159706) and completed May 19, 2009. This is the the first time a 4P Opteron system has exceeded 1-tile-per-core to achieve the highest composite score and bests the previous high score – a Dell R905 – by 1% and 1 tile (102 virtual machines total).

Dell’s M905 was fitted with 128GB of PC2-5300 (DDR2/667) registered ECC memory, 4 on-board Broadcom NetXtreme 1Gbps, and a QLogic QLE2462 FC adapter for virtual machine storage. Three Dell/EMC CX3-4of’s were used including 6 enclosures with 15 disks per enclosure to deliver 18 LUNx (6 per enclosure, 90 physical disks total.)

h1

First 48-core VMmark Appears

June 18, 2009

Following in the footsteps of the first 12-core VMmark comes the current champion at 33.85@24 tiles using 48-cores – and, despite the timing, it is not an Istanbul server. In fact, today’s leader is the IBM System x3950 M2 running 8, 6-core Intel Xeon MP “Dunnington” X7460 processors with 256GB DDR2/667 RAM (5.3GB/core).

This score edges-out the previous champion – the HP ProLiant DL785 G5 with 8, 4-core Opteron 8393SE processors – which reigned at 31.56@21 tiles. In contrast to the 4-socket, 24-core IBM System x3850 M2 Xeon leading the 24-core category, this doubling of socket/core count resulted in only a 50% increase in capacity. This scaling inefficiency is less typical in 2P-to-4P transition but seems to plague the 4p-to-8P segment.

“The x3950 M2 is based on the fourth generation of IBM Enterprise X-Architecture®, and is designed to deliver innovation with enhanced reliability and availability features that enable optimal performance for databases, enterprise applications and virtualized environments.”

IBM News Blurb

“I’m really looking forward to even more virtualization benchmarks which are coming very soon.”

– Elisabeth Stahl, IBM Benchmarking and Systems Performance Blog

Looking at the virtualization notes we discover what it takes to keep 48-cores fed to achieve such a benchmark:

  • 4-QLogic QLE2462 HBA’s (Dual-port, 4-Gbps FC)
  • 1-IBM DS4800 with 4GB cache
    • 19 EXP 810 storage expansion units for
    • 1.8TB in 49 LUNs
      • 280 15K disks total
  • 21 IBM x336 clients
    • DP 3.2GHz Xeon
    • 3GB RAM
    • Server 2003 R2
  • 2 IBM x335 clients
    • DP Xeon 3.06GHz
    • 2.5GB RAM
    • Server 2003 R2
  • Eight vSwitches
    • 120 ports total
  • 4 Intel PRO 1000PT Dual-port 1Gb Ethernet controllers
    • one per vSwitch

While the Dunnington tops the list by sheer brute force, it’s safe to assume that – given the 32-core Opteron is nipping at its heels – the 48-core Istanbul results will displace it soon (possibly alluded to in Elisabeth Stahl’s “Benchmarking and Performance Blog” reference above). More interestingly, will AMD’s much touted “HT Assist” allow the 8P Istanbul to break the 4P-to-8P “curse” of scaling inefficiency? If not, it would show that much work is needed before the relatively “massive ” core counts of 2010 are upon us.