For an industry, digital data storage, that's seen disruptions every 25 years, we're either overdue for one or The Next Revolution has arrived, but nobody's noticed.
What will the Next Big Thing in Storage look like? What, if anything, will succeed RAID, if it hasn't already? Are their lessons we can learn by examining the last big disruption circa 1990 in Storage: RAID arrays?
The one certain lesson from 1955 and 1990 is that nobody can guess, not even remotely, what Data Storage will look like in 25 years time. Any prediction of the 2040 market will be wildly inaccurate, but we can talk about current market forces and technologies and where the trends point for the next 5 and 10 years.
Thirty Years in I.T. Theories, Ideas, Opinions.... Leveraging knowledge of the past to understand now. @SteveJCbr & stevej.cbr@gmail.com
Showing posts with label flash memory. Show all posts
Showing posts with label flash memory. Show all posts
2014/05/22
2012/01/31
Why Fibre Channel SAN's will be dead in 5 years.
I won't be buying shares in any Fibre Channel-based tech stocks as I think the technology will be dead within 5 years for two reasons:
Another article by Fusion-io, Getting the most out of flash storage provides extra links:
Elsewhere I've written Jim Gray's observation, "Disk is the new Tape" should be "Disk is the new CD".
That is, Enterprise Storage is best suited for "Seek and Stream", not random I/O.
Enterprise Storage Arrays will need to provide in this future:
Ethernet and nothing else.
Whether Layer 2 protocols, like Coraid's ATA-over-Ethernet, or Layer 3 protocols, e.g. the slower, higher overhead but routable iSCSI, dominate is still an open question.
Both have strengths and weaknesses and can be used together very effectively without conflicts to maximise ROI's and minimise Enterprise Storage costs, both CapEx and OpEx.
Ethernet is around 10-100 times cheaper than Fibre Channel to install and configure and requires only a fraction of the Administration and support, because Enterprises already have well resourced and competent Networking teams. Network Engineers are in much better supply than SAN specialists, so wages are more reasonable and availability much, much higher.
Their competency/capability is also much better able to be assessed by technical managers when both hiring and firing.
As well, Ethernet has a current growth path of 40Gbps and 100Gbps with 10Gbps widely available now for servers.
Fibre Channel may improve sometime in the future to 12Gbps, but that's an uncertain roadmap and with a global market in 10,000's vs millions for ethernet, the cost differential will only grow.
Fibre Channel has become a very poor choice when bandwidth is the primary "figure of merit".
In access latency, PCI-based Flash Memory, such as Fusion-io's, will always beat SAN-based Storage Arrays by a rather large margin.
It's there in the physics and unbeatable...
All the interfaces, line delays, buffering and switching - out and back on a SAN - means if even the Storage Array SSD's had zero latency, it would be many-times slower.
This Q and A with Matt Young of Fusion-io on "Making Flash Fast", says it well:
Declaration of Interest:
- Enterprise Storage Arrays will stop being used for high-intensity random I/O, instead being used for "seek and stream", and
- PCI Flash or SCM Storage will become the low-latency "Tier 0" Storage of choice because of speed (latency), cost and simplicity.
Another article by Fusion-io, Getting the most out of flash storage provides extra links:
- Flash storage moves closer to CPUs: STEC is moving into PCI Flash cards.
- EMC: Flash could spell doom for Fibre Channel but don't talk about PCI Flash challenges.
- Flash storage in post-PC devices advances, general background on Flash memory.
Elsewhere I've written Jim Gray's observation, "Disk is the new Tape" should be "Disk is the new CD".
That is, Enterprise Storage is best suited for "Seek and Stream", not random I/O.
Enterprise Storage Arrays will need to provide in this future:
- reliable persistent storage and archives,
- high capacity and best $ per MB, and
- high bandwidth streaming IO.
- and if you were being honest, vendor-neutral management and protocols, in-place upgrades, any-time snapshots/backups and flexible zero-downtime expansion and reconfiguration.
The current need for "fork-lift upgrades" and vendor incompatibility are disadvantageous to customers.
Ethernet and nothing else.
Whether Layer 2 protocols, like Coraid's ATA-over-Ethernet, or Layer 3 protocols, e.g. the slower, higher overhead but routable iSCSI, dominate is still an open question.
Both have strengths and weaknesses and can be used together very effectively without conflicts to maximise ROI's and minimise Enterprise Storage costs, both CapEx and OpEx.
Ethernet is around 10-100 times cheaper than Fibre Channel to install and configure and requires only a fraction of the Administration and support, because Enterprises already have well resourced and competent Networking teams. Network Engineers are in much better supply than SAN specialists, so wages are more reasonable and availability much, much higher.
Their competency/capability is also much better able to be assessed by technical managers when both hiring and firing.
As well, Ethernet has a current growth path of 40Gbps and 100Gbps with 10Gbps widely available now for servers.
Fibre Channel may improve sometime in the future to 12Gbps, but that's an uncertain roadmap and with a global market in 10,000's vs millions for ethernet, the cost differential will only grow.
Fibre Channel has become a very poor choice when bandwidth is the primary "figure of merit".
In access latency, PCI-based Flash Memory, such as Fusion-io's, will always beat SAN-based Storage Arrays by a rather large margin.
It's there in the physics and unbeatable...
All the interfaces, line delays, buffering and switching - out and back on a SAN - means if even the Storage Array SSD's had zero latency, it would be many-times slower.
This Q and A with Matt Young of Fusion-io on "Making Flash Fast", says it well:
Q: How does your ioMemory technology differ from Solid State Disks? And how does it compare performance wise?and
A: Solid State Disks or SSDs are used to store data with the intention of constant use – similar to that of a hard drive. These SSDs generally use disk-based protocols that introduce unnecessary latency into the system. Fusion’s ioMemory technology differs in that it doesn’t act as a hard drive. It performs as an extension of the memory hierarchy for servers. This means that they provide a tighter integration with host systems and applications, helping you to work more productively.
Fusion-io products offer the industry’s lowest latencies, which maximise performance and scalability, while delivering enterprise reliability.
Q: Can you provide some typical I/O performance figures for ioMemory compared to DRAM, and solid state disk?
A: With some generalisation, the order of memory, fastest first is as follows,
There are of course a number of factors that need to be considered in these times such as payload size, load, etc, however, in simple terms with all things equal and a well-designed product, latency is ultimately affected by the distance data must travel to get to where it is useful.
- DRAM with 100-300 nanosecond access,
- ioMemory with 15 microsecond access,
- NAND appliances with around 500 microsecond access and
- then SSD’s with around 1ms [1000 microsecond] access. [66 times slower...]
So, the closer your technology resides in relation to the CPU the better the response time.
That’s why that even though two products may use the same NAND chips and be connected on the PCI Express bus, you see markedly different latency characteristics.finally:
Q: As well as I/O performance, what other attributes of ioMemory are finding favour among customers?
A: One of the things that our customers tell us provides a major benefit in addition to performance is the reliability of ioMemory and the cost savings generated from implementing Fusion-io solutions. Fusion-io products are uniquely reliable enough to be offered by all major OEM manufacturers, including Dell, HP and IBM.
[snip]
Finally, many customers tell us that they save a lot of money on CapEx and OpEx, since ioMemory takes so much less power, cooling and real estate than traditional, scaled-out storage infrastructures.
Declaration of Interest:
I have no shares or other financial interest in Fusion-io or any companies or their competitors mentioned in this piece.(signed) Steve Jenkin, 31-Jan-2012.
I am not employed now, nor have ever been, by Fusion-io or any of its related/associated entities.
I receive no remuneration for writing these opinions/analyses.
How they got to 1 Billion IO per second.
Is Fusion-io's demonstration of "1 Billion IO operations per second" the same sort of game-changer that the 1987/8 RAID paper by Patterson, Katz and Gibson was?
Within 5 years all "Single Large Expensive Disks" (SLED's) were out of production, will we see Flash disks in Storage Arrays and low-latency SAN's out-of-production by 2017?
A more interesting "real world" demo by Fusion-io in early 2012 was loading MS-SQL in 16 virtual machines running Windows 2008 under VMware. They achieved a 2.8-fold improvement in throughput with a possible (unstated) 5-10 fold access-time improvement.
Updated 16:00, 31-Jan-2012. A little more interpretation of the demo descriptions and detailed PCI bus analysis.
Fusion used a total of 64 cards in 8 servers running against a "custom load generator", or 16 million IO/sec per card.
There are two immediate problems:
The card spec sheet also quotes throughput of 3GB/sec (24Gbps) for large read and 2.5GB/sec (20Gbps) for large writes (1MB I/O).
There was no information on the read/write ratio of the load, nor if they were random or sequential.
From a piece on Read Caching also speeding up writes by 3-10 fold, Fusion show they are savvy systems designers, they use a 70:30 (R:W) workload as representative of real-world loads.
That nothing was said about the workload suggests it may have pure read or pure write - whichever was faster with ACM under Linux. If the cards' ACM performance tracks the quoted specs (via a block mode interface), this would be pure sequential write.
The workload must have been 50:50 to allow full utilisation of the single shared PCI bus on each system, otherwise the bus would've saturated.
Also, as this is as much a demonstration of ACM, integrated with the Virtual Memory system to cause "page faults", the transfers to/from the Fusion cards were probably in whole VM pages. The VM page size isn't stated.
In Linux pages default to 4KB or 8KB, but are configurable to at least 256MB. Again, these are savvy techs with highly competent kernel developers involved, so an optimum page file for the Fusion card architecture, potentially 1-4MB, was chosen for the demo. [Later, 64KB is used for PCI bus calculations.]
The Fusion write does not say how they checked the IO's were correctly written. With 153.6TB total storage and 64GB/sec in the test work load, the tests could've run 2,500 seconds (40min) before filling the cards. Perhaps they read-back the contents and compared that to the generated input, though nothing is said. In the best of all worlds, there would've been a real-time read-back check, i.e. a 50:50 R:W workload.
The 64GB/sec total I/O throughput gives 8GB/sec, or 64Gbps, per server.
The HP ProLiant DL370 servers used in the tests, according to the detailed specs, only support 9 PCI-e v 2 cards, most slots are 4 lane ('x4'), at 4Gbps per lane, bi-directionally. With 8 slots taken by the Fusion-io cards, only one (x16?) slot was available for the network card needed to supplement the 4 1Gbps on-board ethernet ports.
80-100Gbps ethernet capacity would normally be needed to support 64Gbps of IP traffic.
Reading the "datasheet" carefully, including inferring from the diagram which has no external connections between systems and no external load-gen host, an internal load-generator was used, one per host. There may have been some co-ordination between hosts of the load generators, such as partitioning work units. From the datasheet commentary:
PCI Express v2 is ~4Gbps per lane, or 64Gbps for x16, each direction. Potentially just enough if an on-board ToE was used either with a 8-way 10Gbps card, a dual-port 40Gbps card or a single-port 100Gbps card.
From this, we're can't be sure if the load generator was internal to the Linux servers or external.
Even though not used, HP support dual 10Gbps ethernet cards, PCI-e v2 x8, in the DL370 G6, but with a maximum of two per system. This suggests a normal operational limit of the PCI backplane of 20-40Gbps per direction. The aggregate 64Gbps is achievable if split into 32Gbps in each direction.
The Fusion cards "half-length" and are x8 (so will work with x1, x2, and x4 slots as well).
From the DL370 specs, the system has 8 available full-length slots:
The other half-length slot is probably x8.
The per-server load is 64Gbps, spread amongst 8 cards, or 8Gbps/card, which is only 2 lanes (x2).
The per-card bandwidth would be possible if the single shared PCI bus wasn't saturated.
Per direction, the x4 PCI-e lanes would only support 16Gbps and the x8 32Gbps.
The simple average (5 * 16 + 3 * 32)/8, or 22Gbps, is insufficient for the load.
The two x16 slots and x8 slot could support the maximum transfer rate of the Fusion cards, 32Gbps per direction, or an aggregate of 64Gbps. The cards' spec sheet allows 20-24Gbps large (1MB) transfers per card, which with some load-generator tuning, could've resulted in 60Gbps aggregate from just 3 cards.
If the I/O total load, 64Gbps, is split evenly between cards, each card must process an aggregate 8Gbps, with equal read/write loads, or 4Gbps per direction.
If 64KB pages are read/written, then each card will need to process 64K (65,536) pages per second per direction.
The x4 slots, with 16Gbps available in each direction (aggregate of 32Gbps), will transfer 64KB in 15.28 usec.
The x8 slots will transfer a 64KB half that time, 7.63 usec.
The average 64KB transfer time for the mix of cards (5 * x4, 3* x8) in the system is:
The required 32Gbps per direction seems feasible.
This either says the DL370 has multiple PCI-e buses, not mentioned in the spec sheet, or something else happened.
Within 5 years all "Single Large Expensive Disks" (SLED's) were out of production, will we see Flash disks in Storage Arrays and low-latency SAN's out-of-production by 2017?
A more interesting "real world" demo by Fusion-io in early 2012 was loading MS-SQL in 16 virtual machines running Windows 2008 under VMware. They achieved a 2.8-fold improvement in throughput with a possible (unstated) 5-10 fold access-time improvement.
Updated 16:00, 31-Jan-2012. A little more interpretation of the demo descriptions and detailed PCI bus analysis.
Fusion used a total of 64 cards in 8 servers running against a "custom load generator", or 16 million IO/sec per card.
There are two immediate problems:
- How did they get the IO results off the servers? Presumably ethernet and TCP IP. [No, internal load generator, no external I/O.]
- The spces on those cards (2.4TB ioDrive2 Duo) only rate them for 0.93M or 0.86 sequential IO/sec (write, read) with 512 byte, a 16-fold shortfall.
The card spec sheet also quotes throughput of 3GB/sec (24Gbps) for large read and 2.5GB/sec (20Gbps) for large writes (1MB I/O).
There was no information on the read/write ratio of the load, nor if they were random or sequential.
From a piece on Read Caching also speeding up writes by 3-10 fold, Fusion show they are savvy systems designers, they use a 70:30 (R:W) workload as representative of real-world loads.
That nothing was said about the workload suggests it may have pure read or pure write - whichever was faster with ACM under Linux. If the cards' ACM performance tracks the quoted specs (via a block mode interface), this would be pure sequential write.
The workload must have been 50:50 to allow full utilisation of the single shared PCI bus on each system, otherwise the bus would've saturated.
Also, as this is as much a demonstration of ACM, integrated with the Virtual Memory system to cause "page faults", the transfers to/from the Fusion cards were probably in whole VM pages. The VM page size isn't stated.
In Linux pages default to 4KB or 8KB, but are configurable to at least 256MB. Again, these are savvy techs with highly competent kernel developers involved, so an optimum page file for the Fusion card architecture, potentially 1-4MB, was chosen for the demo. [Later, 64KB is used for PCI bus calculations.]
The 64GB/sec total I/O throughput gives 8GB/sec, or 64Gbps, per server.
The HP ProLiant DL370 servers used in the tests, according to the detailed specs, only support 9 PCI-e v 2 cards, most slots are 4 lane ('x4'), at 4Gbps per lane, bi-directionally.
Reading the "datasheet" carefully, including inferring from the diagram which has no external connections between systems and no external load-gen host, an internal load-generator was used, one per host. There may have been some co-ordination between hosts of the load generators, such as partitioning work units. From the datasheet commentary:
Custom load generator that exercises memory-mapped I/O at a rate of approximately 125 million operations per second on each server. Each operation is a 64-byte packet.A little more information is available in a blog entry.
PCI Express v2 is ~4Gbps per lane, or 64Gbps for x16, each direction.
Even though not used, HP support dual 10Gbps ethernet cards, PCI-e v2 x8, in the DL370 G6, but with a maximum of two per system. This suggests a normal operational limit of the PCI backplane of 20-40Gbps per direction. The aggregate 64Gbps is achievable if split into 32Gbps in each direction.
From the DL370 specs, the system has 8 available full-length slots:
- 2 * x16,
- 1 * x8 and
- 5 * x4.
The per-server load is 64Gbps, spread amongst 8 cards, or 8Gbps/card, which is only 2 lanes (x2).
The per-card bandwidth would be possible if the single shared PCI bus wasn't saturated.
Per direction, the x4 PCI-e lanes would only support 16Gbps and the x8 32Gbps.
The simple average (5 * 16 + 3 * 32)/8, or 22Gbps, is insufficient for the load.
The two x16 slots and x8 slot could support the maximum transfer rate of the Fusion cards, 32Gbps per direction, or an aggregate of 64Gbps. The cards' spec sheet allows 20-24Gbps large (1MB) transfers per card, which with some load-generator tuning, could've resulted in 60Gbps aggregate from just 3 cards.
If the I/O total load, 64Gbps, is split evenly between cards, each card must process an aggregate 8Gbps, with equal read/write loads, or 4Gbps per direction.
If 64KB pages are read/written, then each card will need to process 64K (65,536) pages per second per direction.
The x4 slots, with 16Gbps available in each direction (aggregate of 32Gbps), will transfer 64KB in 15.28 usec.
The x8 slots will transfer a 64KB half that time, 7.63 usec.
The average 64KB transfer time for the mix of cards (5 * x4, 3* x8) in the system is:
( (5 * 15.28) + (3 * 7.63)) / 8, or 12.40 usec,or 80,659 64KB pages per second per direction, leaving 25% headroom for other traffic and bus controller overheads.
The required 32Gbps per direction seems feasible.
2011/09/14
A new inflection point? Definitive Commodity Server Organisation/Design Rules
Summary:
For the delivery of general purpose and wide-scale Compute/Internet Services there now seems to be a definitive hardware organisation for servers, typified by the E-bay "pod" contract.
For decades there have been well documented "Design Rules" for producing Silicon devices using specific technologies/fabrication techniques. This is an attempt to capture some rules for current server farms. [Update 06-Nov-11: "Design Rules" are important: Patterson in a Sept. 1995 Scientific American article notes that the adoption of a quantitative design approach in the 1980's led to an improvement in microprocessor speedup from 35%pa to 55%pa. After a decade, processors were 3 times faster than forecast.]
Commodity Servers have exactly three possible CPU configurations, based on "scale-up" factors:
[Update 06-Nov-11: Because Oracle insists some feature sets must run on raw hardware. Sometimes vendors won't support your (preferred) VM solution.]
VM products are close to free and offer incontestable Admin and Management advantages, like 'teleportation' or live-migration of running instances and local storage.
There is a special non-VM case: cloned physical servers. This is how I'd run a mid-sized or large web-farm.
This requires careful design, a substantial toolset, competent Admins and a resilient Network design. Layer 4-7 switches are mandatory in this environment.
There are 3 system components of interest:
Consequentially, "Fibre Channel over Ethernet" with its inherent contradictions and problems, is unnecessary.
Designing individual service configurations can be broken down into steps:
As a professional, you're looking to provide "bang-for-buck" for someone else who's writing the cheques. Over-dimensioning is as much a 'sin' as running out of capacity. Nobody ever got fired for spending just enough, hence maximising profits.
Getting it right as often as possible is the central professional engineering problem.
Followed by, limiting the impact of Faults, Failures and Errors - including under-capacity.
The quintessential advantage to professionals in developing standard, reproducible designs is the flexibility to respond to unanticipated load/demands and the speed with which new equipment can be brought on-line, and the converse, retired and removed.
Security architectures and choice of O/S + Cloud management software is outside the scope of this piece.
There are many multi-processing architectures, each best suited to particular workloads.
They are outside the scope of this piece, but locally attached GPU's are about to become standard options.
Most servers will acquire what were known as vector processors and applications using this capacity will start to become common. This trend may need their own Design Rule(s).
Different, though potentially similar design rules apply for small to mid-size Beowulf clusters, depending on their workload and cost constraints.
Large-scale or high-performance compute clusters or storage farms, such as the IBM 120 Petabyte system, need careful design by experienced specialists. With any technology, "pushing the envelope" requires special attention by the best people you have, to even have a chance of success.
Not unsurprisingly, this organisation looks a lot like the current fad, "Cloud Computing" and the last fad, "Services Oriented Architecture".
Google and Amazon dominated their industry segments partly because they figured out the technical side of their business early on. They understood how to design and deploy datacentres suitable for their workload, how to manage Performance and balance Capacity and Cost.
Their "workloads", and hence server designs, are very different:
For the delivery of general purpose and wide-scale Compute/Internet Services there now seems to be a definitive hardware organisation for servers, typified by the E-bay "pod" contract.
For decades there have been well documented "Design Rules" for producing Silicon devices using specific technologies/fabrication techniques. This is an attempt to capture some rules for current server farms. [Update 06-Nov-11: "Design Rules" are important: Patterson in a Sept. 1995 Scientific American article notes that the adoption of a quantitative design approach in the 1980's led to an improvement in microprocessor speedup from 35%pa to 55%pa. After a decade, processors were 3 times faster than forecast.]
Commodity Servers have exactly three possible CPU configurations, based on "scale-up" factors:
- single CPU, with no coupling/coherency between App instances. e.g. pure static web-server.
- dual CPU, with moderate coupling/coherency. e.g. web-servers with dynamic content from local databases. [LAMP-style].
- multi-CPU, with high coupling/coherency. e.g. "Enterprise" databases with complex queries.
[Update 06-Nov-11: Because Oracle insists some feature sets must run on raw hardware. Sometimes vendors won't support your (preferred) VM solution.]
VM products are close to free and offer incontestable Admin and Management advantages, like 'teleportation' or live-migration of running instances and local storage.
There is a special non-VM case: cloned physical servers. This is how I'd run a mid-sized or large web-farm.
This requires careful design, a substantial toolset, competent Admins and a resilient Network design. Layer 4-7 switches are mandatory in this environment.
There are 3 system components of interest:
- The base Platform: CPU, RAM, motherboard, interfaces, etc
- Local high-speed persistent storage. i.e. SSD's in a RAID configuration.
- Large-scale common storage. Network attached storage with filesystem, not block-level, access.
Consequentially, "Fibre Channel over Ethernet" with its inherent contradictions and problems, is unnecessary.
Designing individual service configurations can be broken down into steps:
- select the appropriate CPU config per service component
- specify the size/performance of local SSD per CPU-type.
- architect the supporting network(s)
- specify common network storage elements and rate of storage consumption/growth.
As a professional, you're looking to provide "bang-for-buck" for someone else who's writing the cheques. Over-dimensioning is as much a 'sin' as running out of capacity. Nobody ever got fired for spending just enough, hence maximising profits.
Getting it right as often as possible is the central professional engineering problem.
Followed by, limiting the impact of Faults, Failures and Errors - including under-capacity.
The quintessential advantage to professionals in developing standard, reproducible designs is the flexibility to respond to unanticipated load/demands and the speed with which new equipment can be brought on-line, and the converse, retired and removed.
Security architectures and choice of O/S + Cloud management software is outside the scope of this piece.
There are many multi-processing architectures, each best suited to particular workloads.
They are outside the scope of this piece, but locally attached GPU's are about to become standard options.
Most servers will acquire what were known as vector processors and applications using this capacity will start to become common. This trend may need their own Design Rule(s).
Different, though potentially similar design rules apply for small to mid-size Beowulf clusters, depending on their workload and cost constraints.
Large-scale or high-performance compute clusters or storage farms, such as the IBM 120 Petabyte system, need careful design by experienced specialists. With any technology, "pushing the envelope" requires special attention by the best people you have, to even have a chance of success.
Not unsurprisingly, this organisation looks a lot like the current fad, "Cloud Computing" and the last fad, "Services Oriented Architecture".
Google and Amazon dominated their industry segments partly because they figured out the technical side of their business early on. They understood how to design and deploy datacentres suitable for their workload, how to manage Performance and balance Capacity and Cost.
Their "workloads", and hence server designs, are very different:
- Google serves pure web-pages, with almost no coupling/communication between servers.
- Amazon has front-end web-servers is backed by complex database systems.
2010/05/03
A Good Question: When will Computer Design 'stabilise'?
The other night I was talking to my non-Geek friend about computers and he formulated what I thought was A Good Question:
It comes with a 5 year warranty, which leads to the obvious question:
The server market has already fractioned into "budget", "value" and "premium" species.
The desktop/laptop market continues to redefine itself - and more 'other' devices arise. The 100M+ iPhones, in particular, already out there demonstrate this.
There's a new major step in server evolution just breaking:
An interesting side question:
How will Near-Zero-Latency local storage impact system 'performance', both response times (a.k.a. latency) and throughput.
I conjecture that both system latency and throughput will improve markedly, possibly super-linearly, because one of the bug-bears of Operating Systems, the context switch, will be removed. Systems have to expend significant effort/overhead in 'saving their place', deciding what to do next, then when the data is finally ready/available, to stop what they were doing and start again where they left off.
The new processing model, especially for multi-core CPU's, will be:
It would seem that Operating Systems might benefit from significant redesign to exploit this effect, in much the same way that RAM is now large and cheap enough that system 'swap space' is now either an anachronism or unused.
The evolution of USB flash drives saw prices/Gb halving every year. I've recently seen 4Gb SDHC cards at the supermarket for ~$15, whereas in 2008, I paid ~$60 for USB 4Gb.
Rough server pricing for RAM in 2010 is A$65/Gb ±$15.
List prices by Tier 1/2 vendors for 64Gb SSD is $750-$1000 (around 2-4 times cheaper from 'white box' suppliers).
I've seen this firmware limited to 50Gb to improve performance and reliability comparable to current production HDD specs.
This is $12-$20/Gb, depending on what base size and prices used.
Disk drives are ~A$125 for 7200rpm SATA and $275-$450 for 15K SAS drives.
With 2.5" drives priced in-between.
Ie. $0.125/Gb for 'big slow' disks and $1 per GB for fast SAS disks.
Roll forward 5 years to 2015 and 'SSD' might've doubled in size three times, plus seen the unit price drop. Hard disks will likely follow the same trend of 2-3 doublings.
Say SSD 400Gb for $300: $1.25/Gb
2.5" drives might be up to 2-4Tb in 2015 (from 500Gb in 2010) and cost $200: $0.05-0.10/Gb
RAM might be down to $15-$30/Gb.
A caveat with disk storage pricing: 10 years ago RAID 5 became necessary for production servers to avoid permanent data loss.
We've now passed another event horizon: Dual-parity, as a minimum, is required on production RAID sets.
On production servers, price of storage has to factor in the multiple overheads of building high-reliability storage (redundant {disks, controllers, connections}, parity and hot-swap disks and even fully mirrored RAID volumes plus software, licenses and their Operations, Admin and Maintenance) from unreliable parts. A problem solved by electronics engineers 50+ years ago with N+1 redundancy.
Multiple Parity is now needed because in the time taken to recreate a failed drive, there's a significant chance of a second drive failure and total data loss. [Something NetApp has been pointing out and addressing for some years.] The reason for this is simple: the time to read/write a whole drive has steadily increased since ~1980. Recording density (bits per inch) times areal density (tracks per inch) have increased faster than read/write speeds, roughly multiplying recording density times rotational speed.
Which makes running triple-mirrors a much easier entry point, or some bright spark has to invent a cheap-and-cheerful N-way data replication system. Like a general use Google File System.
Another issue is that current SSD offerings don't impress me.
They make great local disk or non-volatile buffers in storage array, but are not yet, in my opinion, quite ready for 'prime time'.
I'd like to see 2 things changed:
This way the hardware can access flash as large, slow memory and the Operating System can fabricate that into a filesystem if it chooses - plus if it has some knowledge of the on-chip flash memory controller, it can work much better with it. It saves multiple sets of interfaces and protocol conversions.
Direct access flash memory will be always be cheaper and faster than SATA or SAS pseudo-drives.
We would then see following hierarchy of memory in servers:
I can't see a role for Fibre Channel outside storage arrays, and these will go if Infiniband speed and pricing continues to drop. Storage Arrays have used SCSI/SAS drives with internal copper wiring and external Fibre interfaces for a decade or more.
Already the premium network vendors, like CISCO, are selling "Fibre Channel over Ethernet" switches (FCoE using 10GE).
Nary a tape to be seen. (Hooray!)
Servers should tend to be 1RU either full-width or half-width, though there will still be 3-4 styles of servers:
Being normally powered down, you'd expect extended lifetimes for disks and electronics.
But they'll need regular (3-6-12 months) read/check/rewrite cycling or the data will degrade and be permanently lost. Random 'bit-flipping' due to thermal activity, cosmic rays/particles and stray magnetic fields is the price we pay for very high density on magnetic media.
Which is easy to do if they are kept in a remote access device, not unlike "tape robots" of old.
Keeping archival storage "on a shelf" implies manual processes for data checking/refresh, and that is problematic to say the least.
3-5 2.5" drives will make a nice 'brick' for these removable backup packs.
Hopefully commodity vendors like Vantec will start selling multiple-interface RAID devices in the near future. Using current commodity interfaces should ensure they are readable at least a decade into the future. I'm not a fan of hardware RAID controllers in this application because if it breaks, you need to find a replacement - which may be impossible at a future date. (fails 'single point of failure' test).
Which presents another question using a software RAID and filesystem layout: Will it still be available in your O/S of the future?
You're keeping copies of your applications, O/S, licences and hardware to recover/access archived data, aren't you? So this won't be a question... If you don't intend to keep the environment and infrastructure necessary to access archived data, you need to rethink what you're doing.
These enclosures won't be expensive, but shan't be cheap and cheerful:
If it is a valuable asset, potentially irreplaceable, then you must be prepared to pay for its upkeep in time, space and dollars. Just like packing old files into archive boxes and shipping them to a safe off-site facility cost money, it isn't over once they are out of your sight.
Electronic storage is mostly cheaper than paper, but it isn't free and comes with its own limits and problems.
Summary:
Other Reading:
For a definitive theoretical treatment of aspects of storage hierarchies, Dr. Neil J Gunther, ex-Xerox PARC, now Performance Dynamics, has been writing about "The Virtualization Spectrum" for some time.
Footnote 1:
Is this idea of multi-speed memory (small/fast and big/slow) new or original?
No: Seymour Cray, the designer of the world's fastest computers for ~2 decades, based his designs on it. It appears to me to be a old idea whose time has come again.
From a 1995 interview with the Smithsonian:
Footnote 2:
The notion of "all files on the network" and invisible multi-level caches was built in 1990 at Bell Labs in their Unix successor, "Plan 9" (named for one of the worst movies of all time).
Wikipedia has a useful intro/commentary, though the original on-line docs are pretty accessible.
Ken Thompson and co built Plan 9 around 3 elements:
They also pioneered permanent point-in-time archives on disk in something appearing to the user as similar to NetApp's 'snapshots' (though they didn't replicate inode tables and super-blocks).
My observations in this piece can be paraphrased as:
When will they stop changing??This was in reaction to me talking about my experience in suggesting a Network Appliance, a high-end Enterprise Storage device, as shared storage for a website used by a small research group.
It comes with a 5 year warranty, which leads to the obvious question:
will it be useful, relevant or 'what we usually do' in 5 years?I think most of the elements in current systems are here to stay, at least for the evolution of Silicon/Magnetic recording. We are staring at 'the final countdown', i.e. hitting physical limits of these technologies, not necessarily their design limits. Engineers can be very clever.
The server market has already fractioned into "budget", "value" and "premium" species.
The desktop/laptop market continues to redefine itself - and more 'other' devices arise. The 100M+ iPhones, in particular, already out there demonstrate this.
There's a new major step in server evolution just breaking:
Flash memory for large-volume working and/or persistent storage.This implies a major re-organisation of even low-end server installations:
What now may be called internal or local disk.
Fast local storage and large slow network storage - shared and reliable.When the working set of Application data in databases and/or files will fit on (affordable) local flash memory, response times improve dramatically because all that latency is removed. By definition, data outside the working set isn't a rate limiting step, so its latency only slightly affects system response time. However, throughput, the other side of the Performance Coin, has to match or beat that of the local storage, or it will become the system bottleneck.
An interesting side question:
How will Near-Zero-Latency local storage impact system 'performance', both response times (a.k.a. latency) and throughput.
I conjecture that both system latency and throughput will improve markedly, possibly super-linearly, because one of the bug-bears of Operating Systems, the context switch, will be removed. Systems have to expend significant effort/overhead in 'saving their place', deciding what to do next, then when the data is finally ready/available, to stop what they were doing and start again where they left off.
The new processing model, especially for multi-core CPU's, will be:
Allocate a thread to a core, let it run until it finishes and waits for (network) input, or it needs to read/write to the network.Near zero-latency storage removes the need for complex scheduling algorithms and associated queuing. It improves both latency and throughput by removing a bottleneck.
It would seem that Operating Systems might benefit from significant redesign to exploit this effect, in much the same way that RAM is now large and cheap enough that system 'swap space' is now either an anachronism or unused.
The evolution of USB flash drives saw prices/Gb halving every year. I've recently seen 4Gb SDHC cards at the supermarket for ~$15, whereas in 2008, I paid ~$60 for USB 4Gb.
Rough server pricing for RAM in 2010 is A$65/Gb ±$15.
List prices by Tier 1/2 vendors for 64Gb SSD is $750-$1000 (around 2-4 times cheaper from 'white box' suppliers).
I've seen this firmware limited to 50Gb to improve performance and reliability comparable to current production HDD specs.
This is $12-$20/Gb, depending on what base size and prices used.
Disk drives are ~A$125 for 7200rpm SATA and $275-$450 for 15K SAS drives.
With 2.5" drives priced in-between.
Ie. $0.125/Gb for 'big slow' disks and $1 per GB for fast SAS disks.
Roll forward 5 years to 2015 and 'SSD' might've doubled in size three times, plus seen the unit price drop. Hard disks will likely follow the same trend of 2-3 doublings.
Say SSD 400Gb for $300: $1.25/Gb
2.5" drives might be up to 2-4Tb in 2015 (from 500Gb in 2010) and cost $200: $0.05-0.10/Gb
RAM might be down to $15-$30/Gb.
A caveat with disk storage pricing: 10 years ago RAID 5 became necessary for production servers to avoid permanent data loss.
We've now passed another event horizon: Dual-parity, as a minimum, is required on production RAID sets.
On production servers, price of storage has to factor in the multiple overheads of building high-reliability storage (redundant {disks, controllers, connections}, parity and hot-swap disks and even fully mirrored RAID volumes plus software, licenses and their Operations, Admin and Maintenance) from unreliable parts. A problem solved by electronics engineers 50+ years ago with N+1 redundancy.
Multiple Parity is now needed because in the time taken to recreate a failed drive, there's a significant chance of a second drive failure and total data loss. [Something NetApp has been pointing out and addressing for some years.] The reason for this is simple: the time to read/write a whole drive has steadily increased since ~1980. Recording density (bits per inch) times areal density (tracks per inch) have increased faster than read/write speeds, roughly multiplying recording density times rotational speed.
Which makes running triple-mirrors a much easier entry point, or some bright spark has to invent a cheap-and-cheerful N-way data replication system. Like a general use Google File System.
Another issue is that current SSD offerings don't impress me.
They make great local disk or non-volatile buffers in storage array, but are not yet, in my opinion, quite ready for 'prime time'.
I'd like to see 2 things changed:
- RAID-3 organisation with field-replaceable mini-drives. hot-swap preferred.
- PCI, not SAS or SATA connection. I.e. they appear as directly addressable memory.
This way the hardware can access flash as large, slow memory and the Operating System can fabricate that into a filesystem if it chooses - plus if it has some knowledge of the on-chip flash memory controller, it can work much better with it. It saves multiple sets of interfaces and protocol conversions.
Direct access flash memory will be always be cheaper and faster than SATA or SAS pseudo-drives.
We would then see following hierarchy of memory in servers:
- Internal to server
- L1/2/3 cache on-chip
- RAM
- Flash persistent storage
- optional local disk (RAID-dual parity or triple mirrored)
- External and site-local
- network connected storage array, optimised for size, reliability, streaming IO rate and price not IO/sec. Hot swap disks and in-place/live expansion with extra controllers or shelves are taken as a given.
- network connected near-line archival storage (MAID - Massive Array of Idle Disks)
- External and off-site
- off-site snapshots, backups and archives.
Which implies a new type of business similar to Amazon's Storage Cloud.
I can't see a role for Fibre Channel outside storage arrays, and these will go if Infiniband speed and pricing continues to drop. Storage Arrays have used SCSI/SAS drives with internal copper wiring and external Fibre interfaces for a decade or more.
Already the premium network vendors, like CISCO, are selling "Fibre Channel over Ethernet" switches (FCoE using 10GE).
Nary a tape to be seen. (Hooray!)
Servers should tend to be 1RU either full-width or half-width, though there will still be 3-4 styles of servers:
- budget: mostly 1-chip
- value: 1 and 2-chip systems
- lower power value systems: 65W/CPU-chip, not 80-90W.
- premium SMP: fast CPU's, large RAM and many CPU's (90-130W ea)
Being normally powered down, you'd expect extended lifetimes for disks and electronics.
But they'll need regular (3-6-12 months) read/check/rewrite cycling or the data will degrade and be permanently lost. Random 'bit-flipping' due to thermal activity, cosmic rays/particles and stray magnetic fields is the price we pay for very high density on magnetic media.
Which is easy to do if they are kept in a remote access device, not unlike "tape robots" of old.
Keeping archival storage "on a shelf" implies manual processes for data checking/refresh, and that is problematic to say the least.
3-5 2.5" drives will make a nice 'brick' for these removable backup packs.
Hopefully commodity vendors like Vantec will start selling multiple-interface RAID devices in the near future. Using current commodity interfaces should ensure they are readable at least a decade into the future. I'm not a fan of hardware RAID controllers in this application because if it breaks, you need to find a replacement - which may be impossible at a future date. (fails 'single point of failure' test).
Which presents another question using a software RAID and filesystem layout: Will it still be available in your O/S of the future?
You're keeping copies of your applications, O/S, licences and hardware to recover/access archived data, aren't you? So this won't be a question... If you don't intend to keep the environment and infrastructure necessary to access archived data, you need to rethink what you're doing.
These enclosures won't be expensive, but shan't be cheap and cheerful:
Just what is your data worth to you?If it has little value, then why are you spending money on keeping it?
If it is a valuable asset, potentially irreplaceable, then you must be prepared to pay for its upkeep in time, space and dollars. Just like packing old files into archive boxes and shipping them to a safe off-site facility cost money, it isn't over once they are out of your sight.
Electronic storage is mostly cheaper than paper, but it isn't free and comes with its own limits and problems.
Summary:
- SSD's are best suited and positioned as local or internal 'disks', not in storage arrays.
- Flash memory is better presented to an Operating System as directly accessible memory.
- Like disk arrays and RAM, flash memory needs to seamlessly cater for failure of bits and whole devices.
- Hard disks have evolved to need multiple parity drives to keep the risk of total data loss acceptably low in production environments.
- Throughput of storage arrays, not latency, will become their defining performance metric.
New 'figures of merit' will be: - Volumetric: Gb per cubic-inch
- Power: Watts per Gb
- Throughput: Gb per second per read/write-stream
- Bandwidth: Total Gb per second
- Connections: Number simultaneous connections.
- Price: $ per Gb available and $ per Gb/sec per server and total
- Reliability: probability of 1 byte lost per year per Gb
- Archive and Recovery features: snapshots, backups, archives and Mean-Time-to-Restore
- Expansion and Scalability: maximum size (Gb, controllers, units, I/O rate) and incremental pricing
- Off-site and removable storage: RAID-5 disk-packs with multiple interfaces are needed.
- Near Zero-latency storage implies reorganising and simplifying Operating Systems and their scheduling/multi-processing algorithms. Special CPU support may be needed, like for Virtualisation.
- Separating networks {external access, storage/database, admin, backups} becomes mandatory for performance, reliability, scaling and security.
- Pushing large-scale persistent storage onto the network requires a commodity network faster than 1Gbps ethernet. This will either be 10Gbps ethernet or multi-lane 3-6Gbps Infiniband.
What might Desktops look like in 5 years?
Other Reading:
For a definitive theoretical treatment of aspects of storage hierarchies, Dr. Neil J Gunther, ex-Xerox PARC, now Performance Dynamics, has been writing about "The Virtualization Spectrum" for some time.
Footnote 1:
Is this idea of multi-speed memory (small/fast and big/slow) new or original?
No: Seymour Cray, the designer of the world's fastest computers for ~2 decades, based his designs on it. It appears to me to be a old idea whose time has come again.
From a 1995 interview with the Smithsonian:
SC: Memory was the dominant consideration. How to use new memory parts as they appeared at that point in time. There were, as there are today large dynamic memory parts and relatively slow and much faster smaller static parts. The compromise between using those types of memory remains the challenge today to equipment designers. There's a factor of four in terms of memory size between the slower part and the faster part. Its not at all obvious which is the better choice until one talks about specific applications. As you design a machine you're generally not able to talk about specific applications because you don't know enough about how the machine will be used to do that.There is also a great PPT presentation on Seymour Cray by Gordon Bell entitled "A Seymour Cray Perspective", probably written as a tribute after Cray's untimely death in an auto accident.
Footnote 2:
The notion of "all files on the network" and invisible multi-level caches was built in 1990 at Bell Labs in their Unix successor, "Plan 9" (named for one of the worst movies of all time).
Wikipedia has a useful intro/commentary, though the original on-line docs are pretty accessible.
Ken Thompson and co built Plan 9 around 3 elements:
- A single protocol (9P) of around 14 elements (read, write, seek, close, clone, cd, ...)
- The Network connects everything.
- Four types of device: terminals, CPU servers, Storage servers and the Authentication server.
- 1Gb of RAM (more?)
- 100Gb of disk (in an age where 1Gb drives where very large and exotic)
- 1Tb of WORM storage (write-once optical disk. Unheard of in a single device)
They also pioneered permanent point-in-time archives on disk in something appearing to the user as similar to NetApp's 'snapshots' (though they didn't replicate inode tables and super-blocks).
My observations in this piece can be paraphrased as:
- re-embrace Cray's multiple-memory model, and
- embrace commercially the Plan 9 "network storage" model.
2008/03/09
Videos on Flash Memory Cards - II
My friend Mark expanded on my idea of "HD DV being irrelevant" - like phone SIM's, video stores can sell/rent videos on flash cards (like SD) sealed in a credit-card carrier.
The issues are more commercial than technical. 8Gb USB flash memory might hit the A$50 price point this year - and A$30 next year. There is a 'base price' for flash memory - around $10-$15.
This inverts the current cost structure of expensive reader/writer and cheap media. Which is perfect for rental/leasing of media - a refundable 'media deposit' works. An added bonus for content owners is a significant "price barrier" for consumers wanting to make a copy. If a 'stack' of 100 SD cards costs $1500 (vs $100 for DVDs), very few people will throw these around 'like candy'.
Mark's comments:
The issues are more commercial than technical. 8Gb USB flash memory might hit the A$50 price point this year - and A$30 next year. There is a 'base price' for flash memory - around $10-$15.
This inverts the current cost structure of expensive reader/writer and cheap media. Which is perfect for rental/leasing of media - a refundable 'media deposit' works. An added bonus for content owners is a significant "price barrier" for consumers wanting to make a copy. If a 'stack' of 100 SD cards costs $1500 (vs $100 for DVDs), very few people will throw these around 'like candy'.
Mark's comments:
Y'know, the more I think of it, the more the SD-embedded-in-a-credit-card has a lot of appeal when the availability and price point for 8Gb SDs is right. It makes it easy to print a picture, title and credits/notices etc on the 'credit card' - something big enough to be readable and a convenient display format and, as you say, nicely wallet-sized. Snap off the SD and you've agreed to the conditions etc, plus the media is now obviously 'used'.
It's a useful format for other distributions too - games, software, etc (Comes to mind that SAS media still comes on literally dozens of CDs in a cardboard box the size of a couple of shoe boxes).
My complete collection of "Buffy" would come in something the size of a can of SPAM or smaller, rather than something the size of a couple of house bricks for the DVD version, or something still the size of a regular paperback for the Blu-Ray version. For collectors of such things, the difference between having many bookshelves taken up by the complete set ofVs a small box of credit card (or smaller) sized objects is significant. The ability to legally re-burn or replace and re-burn the media when it fails is critical though.
SJ: Because of the per-copy encoding to a 'key', stealing expensive collections isn't useful, unless the key is also taken. So those 'keys' have to be something you don't leave in the Video player.
You've covered the DRM aspects and better alternatives to DRM - which also means that I can burn and sign the media I might produce and distribute myself without needing to involve the likes of Sony or Verisign - although that is possible also - which protects the little producer. Include content in Chrissy and Birthday cards - you've seen those Birthday cards with a CD of songs from your birth year - why not a sample of the movies from that year, plus newsreels etc. Good for things like audio books - whole collections. And if the content on an SD gets destroyed, as long as the media is OK, it would be possible to re-burn it. Most current DVD players now also have SD readers as standard.
Surely someone has thought of it already! Part of the attraction of DVD over storing your library on a 2TB USB disk from Dick Smith is the problem of backups. DVD is perceived, incorrectly, as permanent storage. Though I notice some external USB drives now have built-in RAID 1 or RAID 5, but Joe public doesn't see the need (how come I bought a 2TB drive and I only get 1TB?).
Yeah, I think the proposition that SD or similar will become the ubiquitous preferred standard portable, point-of-sale, recording and backup storage media for photos, movies and music, has some credence. There is something to be said for - "you pick it up in your hand; you buy it; it's yours" - over - "downloading and buying some limited 'right to use' ".
2008/03/06
Who cares about HD DV?
Talking to a friend at lunch today, the topic of "Blu-Ray" vs "HD DV" formats came up...
I think "Blu-Ray" may take the market, but it won't be much of a market.
There are just too many competitors for moving around video files:
His response: "they could package them like SIMs - in a snap-off credit card-sized holder". Which is better than any idea I've had on packaging.
And it fulfills the most important criteria:
Practical problems:
The flash needs a 'fuse' that is broken when the card is freed. Preferably an on-chip use counter that can only be factory reset.
To issue a movie to a customer, the encoding key of the video (if present) would be combined with the users key - and the resulting unique key written on the card. Players need both the card and user key to decode and play the movie.
That same process also tags the card with the current owner.
You lose it, it can come home to you.
Because the content can be locked to a particular ID, the raw content can be stored on disk without the movie studios giving away their birth right.
Summary:
I think 120mm disks are going to follow the floppy disk into the technology graveyard.
They will have certain uses - like posting something on cheap, robust media.
With the convergence of PC displays and Home Theater, the whole "Hi-Def TV" problem is morphing. Blu-Ray - can't wait to not buy one.
I think "Blu-Ray" may take the market, but it won't be much of a market.
There are just too many competitors for moving around video files:
- DVD format disks - still good for 8Gb (dual layer). Drives & media are cheap.
- flash memory - 2008 sees A$50 for 8Gb on USB (less on SD card)
- A$300 for 750-,1000Gb USB hard-drives. Under $1/DVD.
- Internet download. With ADSL 2+ giving 5-10Mbps for many.
His response: "they could package them like SIMs - in a snap-off credit card-sized holder". Which is better than any idea I've had on packaging.
And it fulfills the most important criteria:
fits comfortably in a pocket (now a wallet)
Practical problems:
- How to stop people copying the flash and resealing it?
- Some sort of effective copy-protection system would be good.
- Flagging 'ownership' or usage conditions of a movie. Not so much DRM, but 'this is property of XXX'
The flash needs a 'fuse' that is broken when the card is freed. Preferably an on-chip use counter that can only be factory reset.
To issue a movie to a customer, the encoding key of the video (if present) would be combined with the users key - and the resulting unique key written on the card. Players need both the card and user key to decode and play the movie.
That same process also tags the card with the current owner.
You lose it, it can come home to you.
Because the content can be locked to a particular ID, the raw content can be stored on disk without the movie studios giving away their birth right.
Summary:
I think 120mm disks are going to follow the floppy disk into the technology graveyard.
They will have certain uses - like posting something on cheap, robust media.
With the convergence of PC displays and Home Theater, the whole "Hi-Def TV" problem is morphing. Blu-Ray - can't wait to not buy one.
Subscribe to:
Posts (Atom)