Showing posts with label VMware. Show all posts
Showing posts with label VMware. Show all posts

Monday, January 14, 2019

Homelab refresh

Finally replacing my homelab, for two reasons, consisting of three hosts from 2010 it was ancient, and additionally I lost a drive and my vSAN blew up.

vSphere has finally pulled out the x86 instruction emulation code that allowed really old CPUs to work so while 6.7U1 ran on my 5630L CPUs I couldn't do a clean install (would have had to install 6.5 and upgrade) and nested virtualization was becoming limited by the same thing, which is kind of my killer app for a homelab, on the hosts themselves upgrading is OK, but not being able to instantly have a >6.5 nested host was a pain.

I didn't understand vSAN :)  I'd been running it a long time on unsupported everything (controllers, drives, NICs, you name it) and my early mistakes couldn't be easily fixed as if I tried to reconfigure anything on the fly I got error messages rather than actions.  With money I could have fixed it - by replacing the controllers and buying enough disks for an additional disk group and migrating, or doing something ugly like moving data onto a USB drive or 2 bay NAS...didn't come to that anyhow as I lost so much data there was little point in saving any.

The critical thing I hadn't understood was that erasure coding needs a minimum of four hosts, more if you want to do maintenance, so turning it on in my three host cluster was not smart.  One of my SSDs failed and about half my VMs went with it as they must have had blocks on that disk group that couldn't be pulled from elsewhere.  I daresay I could have recovered many of them but in the lab nothing was critical enough to bother, my greatest pang is for my trusty Windows 7 admin VM...I have been way too cowboy in my lab, which was fine a decade ago, when it was a fraction the size, local to me, and using NFS storage.  These days when I blow it up with a pre-release build that I then find can't be upgraded, or by turning on features for fun before I understand the consequences it's a huge effort to recover.  Nested labs make a lot of sense, where I used to almost enjoy (OK enjoy is overstating it, but I did derive some masochistic pleasure from it and revel in being an expert), the installation pains of the VMware suite, many of those have been reduced (finally) to the point there's no learning in that stage of things.

As with the old cluster I got somewhat carried away building a new one, I similarly received some cast off gear for free and supplemented from eBay and stripping my old systems (only reusing the SSDs). I wanted to grow to four nodes without taking up more space, so when I was gifted a 2015 vintage Supermicro Twin I was happy to purchase a second in order to end up with four identical hosts in 4U of space (replacing 3 X 2U boxes).  This particular model has a SAS controller onboard so I can live with the 2 PCIe slots, reusing my Intel Optane 900p NVMe drives* in one for vSAN cache layer, and installing new Intel X710 10 gig NICs in the other.  (If I'd had a third slot I would've reused my X520's in order to have NICs to pass through to VMs when playing with NSX-T etc.)
The build took a long time as I wanted the firmware on the BIOS/IPMI/SAS controller (now in HBA mode) NICs to all be current - all of which takes a lot of power cycles and messing about, I do see why people purchase vSAN ReadyNodes.  These boxes, 2028TP-DC0R, support current Xeons, I'm using E5-2630L v3, which are not very recent, but cost and power effective and importantly Haswell series so good for some time to come.

The X710's were the biggest time suck, I had two fail on me, not sure if I was unlucky with static, or upgrading their firmware bricked them after a power loss or something.  I would've put the X520s in and been done with them but I only had three and I really wanted the four nodes identical.
I also had second thoughts on RAM, having built out with 128GB per node, deciding longevity would be served better with 192 per.  VMware's stack loves RAM and once I have a pretty complete SDDC running plus a few third party integrations I'd be swapping.

I also turned back on Transparent Page Sharing, enabled nested virtualization, and though I don't thing any of my operating systems support it right now TRIM in vSAN.
I'm now a happy camper, building out a nested lab, the below shows resources consumed by my management layer:



* The Optane are awesomely fast, and have ridiculous endurance too for consumer drives, the 280GB have 336GB inside which supposedly isn't used for traditional over-provisioning, but they must use some of it to help deliver that longevity.  I figure that having the cache tier off the main controller saves the queue in that for destaging to my relatively slow consumer grade SSDs too.  (I had some Enterprise SSDs at one point but they also gave me my only SSD failures, out off warranty of course, where the Samsung Pro consumer drives have been issue free)



Bill of materials:

2 X Supermicro 2028TP-DC0R Twin systems (four nodes) (3008 SAS controller onboard)
8 X Intel Xeon E5-2650Lv3 1.8Ghz 12 core
4 X Intel Optane 900P 280GB PCIe (cache drives, not on vSAN HCL)
4 X Supermicro 64GB SATA-DOM
4 X Intel X710DA Dual 10g SFP+ dual port
4 X Intel 1.6TB S3610 SAS SSD
4 X Samsung 850Pro SATA drives (not on vSAN HCL)
64 X 16GB ECC DDR4 DIMM
2 X HPE 5900 switches (48 gigabit, 4 SFP+, 2XQSFP)

Total, about 20K, but over time and much eBay so very approximate.

P.S. Were I doing this again I'd get the 2028TP-DC0TR, which is exactly the same but with Intel X540 ten gig NICs on board, cost difference now is negligible.  

Tuesday, June 23, 2015

VCDX

My VCDX Journey - Fourth time's a charm!

I'm proud to be able to say I'm now VCDX #197. It seems somewhat of a tradition to write one of these, so despite the count of VCDX's reaching two hundred I'll have a crack. I started writing a blog entry back in 2012 when I originally embarked on this, but never finished it or hit publish.  Through the power and elephantine memory of Google here it is:

What led me here?
I recently received my VCAP-DCA result, appropriately enough, right after I presented a session at VMworld, and unexpectedly passed - I hadn't even answered all the questions before running out of time and knew I'd not scored on a couple.  Little did I know that was normal.  I'd already passed DCD, having taken the beta November 2010, so felt with both, plus a lot of time invested in pre and post sales vSphere consulting, that going for it and trying to make the final VCDX 4 defense in Frankfurt is worthwhile.  These days I'm no longer a consultant, instead part of the VMware alliance team at F5 Networks, which is great, but means it's going to take a long time to develop the depth of both architectural and real world understanding that I think are required of the VCDX - so waiting for version 5 doesn't appeal.  VMware are also going great guns enhancing the products with every release, so there's more to cover with each too, it just gets harder.
Reminds me a lot of the CCIE program back in the day - I took it twice, in 2001 and 2002, right when it contained the kitchen sink of networking:  DLSW, IPX, Token-Ring, Appletalk, VoIP, ATM, in addition to all the IP protocols.  VCDX is very different and to some degree you can self-select your specialties in choosing what to include in your submitted design, but still has that tendancy to get broader and broader, whilst remaining just as deep.  Of course you'll get questioned in the defense on whatever you didn't include too.

Back to present:
I didn't get to Frankfurt, I was too slow at getting it all together, but did submit and defend in Burlington May 2012. I now know that I was very close to passing but didn't, so I defended again in Barcelona in October 2012 and again missed, I suspect by further than the first time.
Roll forward a couple of years, my wife and I had a daughter the following year so VCDX was pushed off the stack for a while. Joining the GCoE pre-sales team at VMware in February of '14 things changed, the team was half VCDX's already and our manager made it clear he'd like the whole team to attain it.
I'd been working on a design already as my old one was vSphere 4 so no longer eligible, and I wanted to incorporate NSX. I was probably two-thirds done with the design and had only outlined the operations and implementation guide. The suggestion was made to do a group submission, with us each playing to our strengths in the sections we write and lots of review, and the pure division of labor, targeting the PEX 2015 defense round.
Working as a group wasn't perfect - we were all in different timezones for a start, and the group decision to go for the Cloud track not ideal for me, but we got it done. We were accepted and February rolled around, I did a horrible job in the defense. I was not expert on the material and nerves ate me up, I barely touched the white board and generally did a poor job of demonstrating the skills of a VCDX. Two of us passed though so I couldn't blame the panel nor the submission.
I was determined to go again as soon as I received the result, and used the experience of the defense to improve the design - there were decisions I couldn't justify because I didn't think they were good and plenty of contradictions and typos that had escaped the ten or so reviews. I'd also labbed lots of vCD in the interim, for me no amount of reading can substitute for hands-on time and researching all the tough questions from the first defense had led me to lots of background I had been missing; expert means expert and I had some holes. The defense was much easier for it, though still nervous I could remember enough to not feel like an idiot, and white board a bunch of stuff too. Still finished feeling I'd failed again but after four times I think that's normal.
My advice after all of this? Go for it, it's worth doing for the knowledge you develop along the way if nothing else. Though I'm a bit of a certification collector (I have worked for channel partners so was paid to) I find vendor certs to be a useful training/development path. I didn't anticipate having to persevere quite so long on this, but the two year break and track change contributed, CCIE took me two attempts too, if something's worth doing it's not going to be easy.
As to the VCDX itself, you need to be an expert on the material, both on your own design and the process of getting to a logical design for a fictional customer. Every decision must be justified by the requirements and constraints! If you're anything like me you need to have enough knowledge of it all that when your performance is impaired by nerves you can still demonstrate enough of it to clear the bar.
Up next? Well as a networking guy I'd rather like to get the VCDX-NV, and every two years after renewing my CCIE (with the Data Center exam last year), I toy with idea of taking another lab...
With respect to Walmart my version of their motto would be 'Always be learning. Always'

Tuesday, May 12, 2015

Useful NSX CLI commands

Useful NSX CLI commands

I'm not going to repeat run of the mill install stuff, but just commands that I've found / people have pointed me to when I've hit issues.  I'll add some API stuff in another post at some point as there's a bunch of stuff not in the CLI at all yet too.

When controllers don't deploy (or deploy then get immediately deleted):
Check disk space on specified datastore
Check /var/log/netcpa.log on the ESX hosts for IP pool allocation issues on controllers
Frequently issues arise because of connectivity NSX Manager to Controllers, and the most common of all:  DNS and NTP issues.

In the Manager CLI there's a handy 'show running-config' these days, which doesn't show a whole lot but will show if you fat fingered it's own network settings.
'show manager log follow' tails the main log file, which aids with all kinds of deployment debugging as the errors can be more verbose than in the GUI.


To troubleshoot MTU issues:
‘esxcli network interface list’ ‘esxcli network nic list’ and ‘ping ++netstack=vxlan x.x.x.x -d -s 1600’  where x.x.x.x is the IP of another hosts VTEP.

vCNS commands that may work:
esxcli network vswitch dvs vmware vslan network mapping list --vds-name=myvds --vxlan-id=5001

esxcli network vswitch dvs vmware vxlan list
esxcli network vswitch dvs vmware vxlan config stats set --level 1
esxcli network vswitch dvs vmware vxlan stats list --vds-name=myvmware

esxcli network vswitch dvs vmware vxlan vmknic multicastgroup list --vds-name=myvds --vlan-id=100

esxcli network vswitch dvs vmware vxlan network stats list --vds-name=myvds --vxlan-id=5001

The dvfilter is the bit that sits between the vmnic and the vswitch and does the packet filtering (and presumably steering in the case of the PANW integration)

summarize-dvfilter - gives back a list of filters present on the host
pktcap-uw --dvfilter $filter-name can then be used to sniff traffic, with --PreDVFilter or --PostDVFilter to help figure out if a rule is not doing what is expected.

pktcap-uw -A    Broader packet capture on ESXi

vsipioctl getfwrules -f $filter_name

ESXi - Controller is TCP/1234

Equivalent of a 'sh cam dy' or sh mac-addr'
net-vdr -b –mac default+edge-1

On Edge, debug packet display interface Nic_0 host_192.168.1.1

show log follow

show service ipsec site

Rene's huge page of links:
http://vcdx133.com/2014/10/05/nsx-link-o-rama/

Logging:
When trying to introduce micro-segmentation to a running environment especially, and for debugging for ever after, not to mention for security auditing, logging is somewhat vital.
NSX has lots of logs all in different places, so redirecting them all to Log Insight / some central location is the way to go, to configure add the log host on NSX manager, and the Edges.
The distributed firewall is distributed :)  So add on every ESXi host:
esxcli system syslog config set --loghost=‘udp://192.168.110.241:514' on every ESXi host
esxcli system syslog reload
esxcli network firewall ruleset set --ruleset-id=syslog --enabled=true
esxcli network firewall refresh

Friday, December 5, 2014

Welcome!

Welcome anyone who has managed to discover this!


I'm currently a Solutions Architect with VMware, formerly a VMware specialist with F5 Networks, and prior to that a channel SE dealing with both plus Cisco and NetApp.
I want to share some of my Evernotes, partially selfishly, so I can get at them from anywhere on the web, but hopefully there's some useful nuggets for others too if Google does it's thing.  I expect them to predominantly be notes that will be updated periodically rather than static posts, as I add commands / tips to them as I go.

This is my first attempt at personal blogging, though I've done some professionally before, while at F5 to share information on the F5 plug-in for vCenter (now vRealize) Orchestrator:

https://devcentral.f5.com/articles/automating-application-delivery-with-big-ip-and-vmware-vcenter-orchestrator
https://devcentral.f5.com/articles/vmware-orchestrator-plug-in-version-20

This is still a subject dear to my heart, and there's now the option of using the VMware Dynamic Types plug-in to talk to the F5 REST API rather than using the, still unsupported, F5 provided plug-in.

I'm now on the team that writes and delivers the internal LiVefire training classes - which are intensive classes on all things Software Defined Data Center, especially the integration of it all rather than the products individually.  As a Cisco / F5 Network head, NSX tends to be my focus, but I touch everything else too on the SDDC side at least.

I'm also a bit of a certification collector, with a bunch of Cisco and VMware ones so far, and will share some of that process too.  CCIE especially has proven valuable to me, and I'm still pursuing VCDX - both also have multiple streams so there's a bunch of stuff to discuss there.  


Usual disclaimer: The views expressed anywhere on this site are strictly mine and not the opinions and views of VMware (or F5 Networks for that matter).