From sacadmin Mon Jan 28 13:31:44 2008
Received: from sac.sfbay.sun.com (localhost [127.0.0.1])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0SLVhsQ022819;
	Mon, 28 Jan 2008 13:31:43 -0800 (PST)
Received: (from carlsonj@localhost)
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8/Submit) id m0SLVh6O022811;
	Mon, 28 Jan 2008 13:31:43 -0800 (PST)
Date: Mon, 28 Jan 2008 13:31:43 -0800 (PST)
From: James Carlson <carlsonj@sac.sfbay.sun.com>
Message-Id: <200801282131.m0SLVh6O022811@sac.sfbay.sun.com>
To: PSARC-record@sac.sfbay.sun.com
Subject: Solaris Bridging [PSARC/2008/055 FastTrack timeout 02/04/2008]
Status: RO
Content-Length: 553


Template Version: @(#)sac_nextcase 1.64 07/13/07 SMI
This information is Copyright 2008 Sun Microsystems
1. Introduction
    1.1. Project/Component Working Name:
	 Solaris Bridging
    1.2. Name of Document Author/Supplier:
	 Author:  James Carlson
    1.3  Date of This Document:
	28 January, 2008
4. Technical Description
    See the case directory for more detail

6. Resources and Schedule
    6.4. Steering Committee requested information
   	6.4.1. Consolidation C-team Name:
		ON
    6.5. ARC review type: FastTrack
    6.6. ARC Exposure: open


From carlsonj@phorcys.east.sun.com Mon Jan 28 13:58:30 2008
Received: from sunmail2sca.sfbay.sun.com (sunmail2sca [129.145.155.234])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0SLwUr9024178
	for <psarc-ext@sac.sfbay.sun.com>; Mon, 28 Jan 2008 13:58:30 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail2sca.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m0SLwQVZ011212
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Mon, 28 Jan 2008 13:58:30 -0800 (PST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVD00B0RJPHMX00@brm-avmta-1.central.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Mon, 28 Jan 2008 14:58:29 -0700 (MST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVD00JSWJPH8N90@brm-avmta-1.central.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Mon,
 28 Jan 2008 14:58:29 -0700 (MST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m0SLwSYf023820	for
 <psarc-ext@sun.com>; Mon, 28 Jan 2008 16:58:28 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m0SLwS06023817; Mon,
 28 Jan 2008 16:58:28 -0500 (EST)
Date: Mon, 28 Jan 2008 16:58:28 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: 2008/055 Solaris Bridging
To: psarc-ext@sun.com
Message-id: <18334.20484.557823.2174@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
Status: RO
Content-Length: 25724

I'm sponsoring this fast-track request for myself.  The timer is set
to 02/04/2008.  The release binding is "Minor" (because we depend on
projects that have Minor binding and do not plan to work on a
back-port) and the interface stability is listed at the end of the
document.

(Of particular interest may be section 5, as that documents
alternatives that were considered and rejected as part of the design.)



This project adds basic Ethernet (layer two) bridging support to
OpenSolaris.  It consists of a Project Private kernel module and
daemon, some Project Private SMF properties, and Committed dladm and
SMF control interfaces.  It is targeted for a Minor release of an
OpenSolaris distribution, though we do not believe that any of the
changes here require Minor binding.

This project assumes that Clearview UV (PSARC 2006/499) will integrate
first.  The terminology and command line design reflects that
assumption.  In particular, Clearview obsoletes the idea of "network
devices" and instead relies on "links" that may themselves be of
varying types.

The bridging protocol referred to in this document is the IEEE
802.1D-1998 "Spanning Tree Protocol," abbreviated in this document as
"STP."  The newer and far more complex "Multiple Spanning Tree
Protocol" (802.1Q-2005; MSTP) is intended to be backward compatible
with STP, and is not part of this project, but may be the subject of a
future project.

This document is large, but we believe that the changes described here
are straightforward and obvious, given the existing system design, and
they've been reviewed by the Clearview and NWAM teams, and thus the
changes are suitable for fast-track treatment.


1.  Administration

    All of the administration of this feature is based on dladm and
    SMF.  The SMF portion is the ability to enable and disable bridge
    instances using the instance URIs described in section 3 below.

1.1 New dladm subcommands

    These commands are patterned after the existing aggregation
    commands in dladm.

    dladm create-bridge [-t] [-R <root-dir>] [-p <priority>]
      [-m <max-age>] [-h <hello-time>] [-d <forward-delay>]
      [-f <force-protocol>] [-l <link>]... <bridge-name>

      This command creates a bridge instance and optionally assigns
      network links to the new bridge.  By default, no bridge
      instances are present, and OpenSolaris will not bridge between
      network links.  See the "add-bridge" subcommand for details on
      link assignment.

      Bridge creation and link assignment require PRIV_SYS_NET_CONFIG.

      In order to bridge between links, you must create at least one
      bridge instance.  Each instance is separate: there is
      intentionally no forwarding connection between bridges.  (Note
      that Crossbow's VNICs may in the future allow virtual inter-
      bridge connections.)

      The <bridge-name> provided is chosen by the administrator and
      arbitrary, but must be a legal SMF service instance name.  For
      purposes of documentation, this is a URI component without
      escape sequences, meaning that the following characters may not
      be present:

	; / ? : @ & = + $ , % < > # "

      including whitespace and ASCII control characters.  The name
      "default" is reserved, as are all names beginning with the
      string "SUNW".  Names with trailing digits are not permitted, in
      order to allow for creation of "observability devices;" see
      section 2 below.

      Because of the use of the observability devices, the names of
      legal bridge instances are further constrained to be a legal
      dlpi(7P) name, which matches:

	[A-Za-z_][A-Za-z0-9_]*[A-Za-z_]

      Options are:

      -t	Create a temporary bridge.

		This will create the bridge object on the running
		system, but the newly created bridge will not survive
		the next reboot.

      -R <root-dir>
		Specify an alternate root directory.

		This allows the configuration of bridge instances in
		alternate roots, as with Live Upgrade and with
		jumpstart installs.  Note that error checking for link
		type isn't possible when administering an alternate
		root.

      -p <priority>
		Specify the Bridge Priority.

		This sets the STP priority value for determining the
		root bridge node in the network.  The default value is
		(per the specification) 32768, and legal values are 0
		(highest priority) to 65535 (lowest priority).

      -m <max-age>
		Specify the maximum age for configuration information.

		This sets the STP Bridge Max Age parameter.
		Information older than this (in seconds) is discarded
		by all bridges in the network if this node is the root
		bridge.  It defaults to 20.0 seconds.  Legal values
		are from 6.0 to 40.0 seconds.  (See the "-d
		<forward-delay>" parameter for additional
		constraints.)

      -h <hello-time>
		Specify the Bridge Hello Time.

		This sets the STP Bridge Hello Time parameter.  If
		this node is the root node, it sends Configuration
		BPDUs at this interval throughout the network.  It
		defaults to 2.0 seconds.  Legal values are from 1.0 to
		10.0 seconds.  (See the "-d <forward-delay> parameter
		for additional constraints.)

      -d <forward-delay>
		Specify the Bridge Forward Delay.

		This sets the STP Bridge Forward Delay parameter.
		This timer is used to sequence the link states when a
		port is enabled anywhere in the network if this node
		is the root bridge.  It defaults to 15.0 seconds.
		Legal values are from 4.0 to 30.0 seconds.

		Bridges must obey the following two constraints:

			2 * (forward_delay - 1.0) >= max_age

			max_age >= 2 * (hello_time + 1.0)

		Any parameter setting that would violate those
		constraints will be treated as an error and cause the
		command to fail with a diagnostic message.

      -f <force-protocol>
		Specify the forced maximum supported protocol.

		This sets the MSTP maximum supported protocol number.
		The default is 3.  The current implementation doesn't
		support RSTP or MSTP, so this currently has no effect.
		However, if the user desires to prevent MSTP from
		being used in the future when implemented, the
		parameter may be set to 0 (STP only) or 2 (allow
		RSTP).

      -l <link>	Add a link to the newly-created bridge.

		This is equivalent to creating the bridge and then
		adding one or more links, as with the "add-bridge"
		option below, except that if any of the links cannot
		be added, then the entire command fails, and the new
		bridge itself isn't created.

    dladm modify-bridge [-t] [-R <root-dir>] [-p <priority>]
      [-m <max-age>] [-h <hello-time>] [-d <forward-delay>]
      [-f <force-protocol>] <bridge-name>

      This subcommand modifies the operational parameters of a given
      bridge instance.  All of the options are the same as for the
      "create-bridge" subcommand above, except that the "-l" option is
      not permitted.  To add links to an existing bridge, use the
      "add-bridge" subcommand below.

      Bridge parameter modification requires PRIV_SYS_NET_CONFIG.

    dladm delete-bridge [-t] [-R <root-dir>] <bridge-name>

      This subcommand deletes a bridge instance.  Unlike the bridge
      creation subcommand, which can add links while creating, it does
      not have the option to remove links during the deletion process.
      The bridge being deleted must not have any attached links.  If
      it does, then an error is returned and no action is taken.

      Bridge deletion requires PRIV_SYS_NET_CONFIG.

      The "-t" and "-R" options are the same as for the
      "create-bridge" subcommand.

    dladm add-bridge [-t] [-R <root-dir>] -l <link> [-l <link>]...
      <bridge-name>

      This subcommand adds one or more links to a bridge instance.  If
      multiple links are specified, and adding any one of them results
      in an error, then no changes are made to the system and the
      command fails.

      Link addition to a bridge requires PRIV_SYS_NET_CONFIG.

      A link may be a member of at most one bridge.  It's an error to
      specify that a link belongs to more than one bridge.  To move a
      link from one bridge instance to another, remove it from the
      current bridge before adding it to the new one.

      The links assigned to a bridge must not themselves be VLANs or
      tunnels.  Only links that would be acceptable as part of an
      aggregation or links that are aggregations themselves may be
      assigned to a bridge.  Other link types will result in error
      messages, and no action taken.  (A future project may provide
      bridging over tunnels using GRE, and over PPP using BCP.  Those
      cases are not part of this project, but nothing this project is
      doing will preclude those cases from the future.)

      In this initial version, the links must also be Ethernet type.
      Bridging is well-defined over a few other media, and there are
      some dodgy ways to make it work on still others, but those cases
      are subjects for a future release.

      When links are added to a bridge, the bridging protocol in use
      (STP) will be notified, and the links will behave as though just
      created.  For STP, this means that the link will be shut down
      and then brought back up using the standard protocol.

      The options are the same as for the "create-bridge" subcommand.

    dladm remove-bridge [-t] [-R <root-dir>] -l <link> [-l <link>]...
      <bridge-name>

      This subcommand removes one or more links from a bridge
      instance.  If multiple links are specified, and removing any one
      of them would result in an error, then none are removed and the
      command fails.

      Link removal from a bridge requires PRIV_SYS_NET_CONFIG.

      When links are removed from a bridge, the bridging protocol
      (STP) is notified, and will likely recalculate a new network
      topology, unless those links were unused due to loop-pruning
      activity by the bridging protocol.

      The options are the same as for the "create-bridge" subcommand.

    dladm show-bridge [-p] [-s [-i <interval>]] [<bridge-name>]

      This subcommand shows the running status of bridges.  When given
      a bridge name, it shows the status of that one bridge.  If no
      bridge name is given, then it shows summary status of all
      bridges on the system.

      Note the lack of a "-R" option here.  It is not possible to list
      bridge configuration information in an alternate root, in
      keeping with the rest of the dladm user interface.  The reason
      for this restriction is to allow the data to be represented in
      SMF, where "writing" to an alternate root is supported by way of
      copying appropriate commands to $ROOT/var/svc/profile/upgrade,
      but "reading" is not feasible because the repository on the
      alternate root may be incompatible with the running system.

1.2 New dladm Link Properties

    "stp"

	This is a boolean property.  It defaults to "true."  When set
	to "false," the link will not use Spanning Tree, and will be
	placed into forwarding mode at all times.  The "false" setting
	is appropriate for point-to-point links connected to end
	nodes.  Only non-VLAN type links have this property.

    "forward"

	This is a boolean property on all links.  It defaults to
	"true."  When set to "false," the VLAN associated with the
	link instance will not forward traffic through the bridge.
	Setting the property to "false" is equivalent to removing the
	VLAN from the "allowed set" for a traditional bridge.

    "default-tag"

	This is a numeric property with range 0 to 4094.  It defaults
	to 1.  It defines the default VLAN ID that's assumed for
	untagged packets sent to and received from this link.  Only
	non-VLAN type links have this property.

    "stp-priority"

	This is a numeric property with range 0 to 255.  It defaults
	to 128.  It corresponds to the STP Port Priority value, which
	is used to determine the preferred root port on a bridge by
	prepending to the port identifier.  Lower numerical values are
	higher priority.

    "stp-cost"

	This is a numeric property with range 1 to 65535; zero is not
	allowed.  It represents the cost for using the link, and
	defaults to (per the standard) 100 for 10Mbps, 19 for 100Mbps,
	4 for 1Gbps, and 2 for 10Gbps.

    "bridge-port"

	This is a read-only numeric property.  It shows the port
	number for the link as seen by the bridge, and is used in
	Spanning Tree messages and network management.

1.3 New Kstats

    Each bridge instance will have a set of statistics, named
    "bridge:<index>:<bridge-name>:<statistic>", where:

	<index>
		Arbitrary instance number assigned by the kernel and
		not necessarily retained across reboot.

	<bridge-name>
		Administrator-specified bridge name.

	<statistic>
		Name of statistic; at least the following:

		learn_source	Number of sources learned
		learn_expire	Number of learnt entries expired
		learn_size	Current count of learnt entries
		forward_direct	Directly forwarded packet count
		forward_unknown	Forwarded with unknown destination
		forward_mbcast	Forwarded multicast/broadcast

    Each link instance will also have new kstats, where the
    <statistic> names will be:

	bridge_sent	Packets forwarded to the link by bridging
	bridge_rcvd	Packets received from the link (and forwarded
			elsewhere) by bridging

    All of these statistics are considered Volatile for now.  The
    existence of the statistics will be documented for users, but with
    warnings that the names and definitions of the statistics may
    change incompatibly.  A future case for the overall RBridges
    project will elevate these in stability.


2.  Packet Observability

    Each bridge instance will be assigned an "observability device,"
    in a manner similar to the DLPI nodes created for "Clearview: IP
    Observability Devices" (PSARC 2006/475).  These nodes will appear
    under the /dev/bridge/ directory, named by the bridge name plus a
    trailing "0".

    The observability node is intended for use with snoop and
    wireshark.  It behaves as a standard Ethernet interface, but does
    not permit the transmission of packets.  All transmitted packets
    are silently dropped.

    The user of this node will get a single unmodified copy of every
    packet handled by the bridge, similar to a "monitoring" port on a
    traditional bridge, and subject to the usual DLPI "promiscuous
    mode" rules.  The user may also filter on VLAN ID by using the
    VLAN PPA hack mechanism: "/dev/bridge/my-bridge1000" selects VLAN
    ID 1 on bridge the instance named "my-bridge".

    The observability node also forms a Project Private control node
    for the kernel, allowing ioctls to a specific bridge instance, and
    will be used by the STP daemon and other (future) bridging
    protocols.

    The dlpi_open(3DLPI) interface will be enhanced with a Committed
    DLPI_BRIDGE flag to allow applications to locate the observability
    nodes by name.


3.  STP Daemon

    Each bridge (created via "dladm create-bridge") is represented as
    an identically-named SMF instance of svc:/network/bridge.  Each
    instance runs a copy of /usr/lib/bridged, which implements the
    Spanning Tree Protocol (STP).  For example, if the user runs:

	# dladm create-bridge my-bridge

    The system will have an SMF service named:

	svc:/network/bridge:my-bridge

    and (per section 2 above) an observability node named:

	/dev/bridge/my-bridge0

    By default, all ports run standard STP.  This is done for safety
    reasons: a bridge that does not run some form of bridging protocol
    (such as STP) can form long-lasting forwarding loops in the
    network.  Because Ethernet has no hop-count or TTL on packets, any
    such loops are fatal to the network.

    When the adminstrator knows that a particular port is not
    connected to another bridge (for example, a direct point-to-point
    connection to a host system), STP can be disabled administratively
    for that port.  Even if all ports on a bridge have STP disabled,
    the STP daemon still runs; this is in case new ports are added,
    and because it is responsible for enabling and disabling
    forwarding on the ports.

    If the SMF service instance for a bridge is disabled, then bridge
    forwarding stops on those ports as the STP daemon is stopped.  If
    the instance is restarted, STP starts from its initial state.

    The bridge daemon runs as UID/GID "daemon" with
    PRIV_SYS_NET_CONFIG in order to access the raw network devices,
    but with most other basic privileges (e.g., PRIV_PROC_FORK and
    PRIV_PROC_EXEC) removed.


3.  VLANs

    In general, administrators will want to have the VLANs they
    configure on the system to be forwarded among all the ports on a
    bridge instance, so this will be the default for VLANs.  When the
    administrator invokes Clearview's "dladm create-vlan", and the
    underlying link is part of a bridge, that command will also enable
    forwarding of the specified VLAN on that bridge link.

    If an administrator wants to configure a VLAN on a link but not
    allow forwarding to or from other links on the bridge, then he
    must take specific action to do so, by disabling forwarding with
    "set-linkprop".

    Clearview UV provides two mechanisms for the creation of VLANs.
    The primary means of configuration is the new "dladm create-vlan"
    subcommand, which automatically enables the VLAN for bridging as
    described above, if the underlying link is configured as part of a
    bridge.

    The second mechanism is a legacy feature called the "PPA hack."
    This allows a user to create a VLAN simply by opening a DLPI
    provider and specifying a VLAN ID number as part of the PPA.  In
    this case, the user may be doing nothing other than snooping on
    that VLAN, so adding the VLAN to the allowed set automatically is
    likely not the right answer.  Thus, we will default forwarding to
    "off" for PPA-hack VLANs.  Administrators with legacy PPA hack
    VLANs will need to reconfigure to use the new Clearview VLANs to
    take full advantage of bridging, and this will be included in the
    documentation.

    In STP, VLANs are ignored.  The bridging protocol computes just
    one loop-free topology and uses that.  Administrators are required
    to configure any "duplicate" links such that when they're
    automatically disabled by STP, the configured VLANs are not
    disconnected.  MSTP is somewhat similar, but allows administrators
    to assign each VLAN to a small number of distinct spanning tree
    "instances," and allows instances within an identically-configured
    "region" to have distinct topologies.  In terms of this project,
    additional bridge and link properties would be required to enable
    MSTP operation.


4.  SMF Properties

    These parameters are all Project Private.  They will not be
    documented, and the documented administrative interface will be
    the dladm command.

4.1 STP SMF

    Property Name		Type		Default
    --------------		----		-------
    config/priority		ushort_t	32768
    config/max-age		ushort_t	5120	(20 seconds)
    config/hello-time		ushort_t	512	(2 seconds)
    config/forward-delay	ushort_t	3840	(15 seconds)
    config/force-protocol	int		3

    All of these properties (and their default values and
    granularities) are defined by the STP and related standards.

    The "force-protocol" parameter is specified to allow for an
    upgrade path.  Users who do not want to see the use of MSTP when
    it is implemented can set this parameter to 0 or 2 (as specified
    in IEEE 802.1Q-2004) to select STP or RSTP as the maximum allowed
    protocol.  In this project, the parameter will have no effect, as
    only STP is implemented.

4.2 Datalink SMF

    Property Name		Type		Default
    --------------		----		-------
    config/stp			boolean		true
    config/forward		boolean		true
    config/bridge		string		""
    config/default-tag		ushort_t	1

    On a Nemo device, legacy device, or aggregation, the link
    parameters are used as above.  The "default-tag" parameter may be
    set to 0 to disable the forwarding of untagged packets to and from
    the port.

    On a VLAN, "stp" and "default-tag" are ignored.  The "forward"
    flag enables forwarding for that VLAN, which is equivalent to
    putting the VLAN into the "allowed set" for the bridge port.
    Setting it to "false" causes the VLAN to be disallowed, which
    means that VLAN-based I/O to the underlying link still operates,
    but no bridge-based forwarding is done.  The "bridge" parameter is
    reserved for use with MSTP, where it will select an instance.


5.  Alternatives

5.1 Using A Separate Command

    An alternative command set design would be to create a new bridge
    control command (bridgeadm), rather than using dladm.

    The main problem with this separation is that the configuration of
    the bridge would end up being split between two different
    utilities in a somewhat incoherent manner.  Why would IEEE 802
    aggregations be part of dladm but IEEE 802 bridges be configured
    elsewhere?

    Parts of the configuration of a bridge (such as the set of allowed
    VLANs and the default VLAN tag for a given link) are naturally
    part of the link configuration, and not a common property of the
    bridge.  The creation of VLANs (logically located "above" links
    and bridges) and regular Ethernet links (logically located "below"
    VLANs and bridges) via dladm while bridging itself is in bridgeadm
    seems like a very strange result.

    We could create a separate bridgeadm, but then we'd likely have to
    deal with the VLAN issues some other way.  Most likely, we would
    end up with either duplicate configuration in bridgeadm or the
    bulk of bridge configuration actually going on in dladm per-link
    properties, and only bridge create/destroy done via bridgeadm.

    In other words, there are several IEEE-specified parameters for
    bridges, but they're rarely of much interest, so that proposed
    utility wouldn't do very much.  The main thing users need to
    manipulate for bridges are the VLANs, and we need to figure out
    how to represent that manipulation.  We choose to equate dladm-
    created-VLAN with bridge-allowed-VLAN because it seems to produce
    the most natural results: there's only one way to "create" or
    "destroy" a VLAN in the system.

    The alternative is to break those apart, and allow users to create
    VLANs for potential use with IP via dladm, and separately assign
    VLANs to bridge ports via bridgeadm, but that runs the very likely
    risk of misconfiguration: either forgetting to enable a bridge
    link for a VLAN while having IP plumbed atop, or thinking that
    destroying the VLAN removes it from the bridge.  Since neither
    scenario seems to be particularly useful, allowing for them
    doesn't seem like a good goal.

    Or, for a really short answer: dladm is the location of all things
    datalinkish, and bridging is (like VLANs and aggregations) a
    datalink function.

5.2 Link Configuration Storage

    Alternative designs for the configuration information include
    having the set of links for a bridge listed as part of the bridge
    configuration, and using non-SMF files for storing configuration.

    The former approach would work, and would have the advantage that
    during start-up of the STP daemon it would be easy to find the
    list of links configured for that instance.  That's a benefit over
    the proposed design in that we will need to iterate over all links
    to get the list needed for a single instance.  However, there are
    two reasons this approach wasn't chosen:

	a. A link may be a member of at most one bridge.  This
	   semantic is easy to enforce with a link property, as
	   there's just one instance of the property, but is hard to
	   enforce across multiple bridges.  We end up needing to scan
	   all bridge instances, and configuration transactions become
	   more complex because two objects need to be changed at one
	   time.

	b. We want to have all configuration parameters for a link to
	   be stored with the link itself.  Having parameters stored
	   elsewhere in the system means that utilities that
	   manipulate links or just display system configuration may
	   end up needing to scan through these other locations in
	   order to make coherent system changes.  (For this project,
	   we would be forced to change the existing Clearview "dladm
	   delete-link" functionality so that it scanned the bridge
	   instances and removed any links found there.  Storing the
	   data with the link instance removes that requirement.)

    Using non-SMF files would also work, and we could make use of the
    Clearview UV "link IDs" to avoid problems inherent with link
    renaming.  However, longer term, the Clearview and NWAM teams are
    refactoring link configuration into SMF.  Having native bridging
    designed for OpenSolaris but not actually integrated with its core
    administrative mechanisms seems like a poor recipe for the future.


6.  Interface Summary

    Interface		Stability		Comments
    ---------		---------		--------
    dladm *-bridge	Committed
    link properties	Committed
    kstats		Volatile		Should be raised later
    /dev/bridge/	Committed		Observability node
    control ioctls	Project Private
    /usr/lib/bridged	Project Private
    /network/bridge	Committed		SMF URI
    config/*		Project Private		SMF properties
    bridge module	Project Private		Kernel bridging module
    DLPI_BRIDGE		Committed		dlpi_open(3DLPI)

From unixconsole@yahoo.com Mon Jan 28 14:19:20 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0SMJJxL024547
	for <psarc-ext@sac.sfbay.sun.com>; Mon, 28 Jan 2008 14:19:20 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id m0SMJHVV025201
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Mon, 28 Jan 2008 22:19:18 GMT
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVD0050LKO5VI00@nwk-avmta-2.sfbay.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Mon, 28 Jan 2008 14:19:17 -0800 (PST)
Received: from brmea-mail-2.sun.com ([192.18.98.43])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVD008TCKO3M6A0@nwk-avmta-2.sfbay.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Mon,
 28 Jan 2008 14:19:15 -0800 (PST)
Received: from relay25.sun.com
 (relay21.sun.com [192.12.251.14] (may be forged))	by brmea-mail-2.sun.com
 (8.13.6+Sun/8.12.9) with ESMTP id m0SM5R50004940	for <psarc-ext@sun.com>; Mon,
 28 Jan 2008 22:19:15 +0000 (GMT)
Received: from mms23es.sun.com ([150.143.232.54] [150.143.232.54])
 by relay25i.sun.com with ESMTP id BT-MMP-1534079 for psarc-ext@sun.com; Mon,
 28 Jan 2008 22:19:14 +0000 (Z)
Received: from relay22.sun.com (relay22.sun.com [192.12.251.34])
 by mms23es.sun.com with ESMTP id BT-MMP-381680 for psarc-ext@sun.com; Mon,
 28 Jan 2008 22:19:13 +0000 (Z)
Received: from web30801.mail.mud.yahoo.com ([68.142.200.144] [68.142.200.144])
 by relay22i.sun.com id BT-MMP-10733504 for psarc-ext@sun.com; Mon,
 28 Jan 2008 22:19:12 +0000 (Z)
Received: (qmail 63924 invoked by uid 60001); Mon, 28 Jan 2008 22:19:12 +0000
Received: from [71.52.72.89] by web30801.mail.mud.yahoo.com via HTTP; Mon,
 28 Jan 2008 14:19:12 -0800 (PST)
Date: Mon, 28 Jan 2008 14:19:12 -0800 (PST)
From: Octave Orgeron <unixconsole@yahoo.com>
Subject: Re: 2008/055 Solaris Bridging
To: James Carlson <james.d.carlson@sun.com>, psarc-ext@sun.com
Message-id: <401216.62773.qm@web30801.mail.mud.yahoo.com>
MIME-version: 1.0
X-Mailer: YahooMailRC/818.31 YahooMailWebService/0.7.162
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
DomainKey-Signature: a=rsa-sha1; q=dns; c=nofws;  s=s1024; d=yahoo.com;
 h=X-YMail-OSG:Received:X-Mailer:Date:From:Subject:To:MIME-Version:Content-Type:Message-ID;
 b=yaszF+UfkfQwotLyfyswyWMaiOaeGgsDX2FZoVCwYcw8zMpQEIRy1XOOUxEB9B2aeR8gG9p6BGwztdNeJEsMao4+7rF5I6kazoxfzwaSvV6ywH3kcBRck6PmkXhoB0B7wye9VYHHNtrY7foAFkhIxeroohQqp2rm9nKyQf5gFQg=;
X-PMX-Version: 5.2.0.264296
X-Brightmail-Tracker: AAAAAA==
X-YMail-OSG: 
 ELJPar4VM1kMDf6CV7vJN1LOUGepRlaPdxHRRcPV53iDL_T_lWSG1a9VVv45nK4Obw55x64Hp5GJnSP04SwoaM3FkjeQpCPg1DjPoEF.edC6FzIzENs-
X-Antispam: No, score=-2.6/5.0, scanned in 0.441sec at (localhost [127.0.0.1])
	by smf-spamd v1.3.1 - http://smfs.sf.net/
Status: RO
Content-Length: 31805

Hi James,

It's good to see this come up for integration. I was wondering how it'll impact virtualization (containers, xen, and ldoms)? Will further work be required or will it work transparently with things like VSW's in LDoms?
 
*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*
Octave J. Orgeron
Solaris Systems Engineer
http://www.opensolaris.org/os/community/sysadmin/
http://unixconsole.blogspot.com
unixconsole@yahoo.com
*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*-*

----- Original Message ----
From: James Carlson <james.d.carlson@sun.com>
To: psarc-ext@sun.com
Sent: Monday, January 28, 2008 3:58:28 PM
Subject: 2008/055 Solaris Bridging

I'm 
sponsoring 
this 
fast-track 
request 
for 
myself.  
The 
timer 
is 
set
to 
02/04/2008.  
The 
release 
binding 
is 
"Minor" 
(because 
we 
depend 
on
projects 
that 
have 
Minor 
binding 
and 
do 
not 
plan 
to 
work 
on 
a
back-port) 
and 
the 
interface 
stability 
is 
listed 
at 
the 
end 
of 
the
document.

(Of 
particular 
interest 
may 
be 
section 
5, 
as 
that 
documents
alternatives 
that 
were 
considered 
and 
rejected 
as 
part 
of 
the 
design.)



This 
project 
adds 
basic 
Ethernet 
(layer 
two) 
bridging 
support 
to
OpenSolaris.  
It 
consists 
of 
a 
Project 
Private 
kernel 
module 
and
daemon, 
some 
Project 
Private 
SMF 
properties, 
and 
Committed 
dladm 
and
SMF 
control 
interfaces.  
It 
is 
targeted 
for 
a 
Minor 
release 
of 
an
OpenSolaris 
distribution, 
though 
we 
do 
not 
believe 
that 
any 
of 
the
changes 
here 
require 
Minor 
binding.

This 
project 
assumes 
that 
Clearview 
UV 
(PSARC 
2006/499) 
will 
integrate
first.  
The 
terminology 
and 
command 
line 
design 
reflects 
that
assumption.  
In 
particular, 
Clearview 
obsoletes 
the 
idea 
of 
"network
devices" 
and 
instead 
relies 
on 
"links" 
that 
may 
themselves 
be 
of
varying 
types.

The 
bridging 
protocol 
referred 
to 
in 
this 
document 
is 
the 
IEEE
802.1D-1998 
"Spanning 
Tree 
Protocol," 
abbreviated 
in 
this 
document 
as
"STP."  
The 
newer 
and 
far 
more 
complex 
"Multiple 
Spanning 
Tree
Protocol" 
(802.1Q-2005; 
MSTP) 
is 
intended 
to 
be 
backward 
compatible
with 
STP, 
and 
is 
not 
part 
of 
this 
project, 
but 
may 
be 
the 
subject 
of 
a
future 
project.

This 
document 
is 
large, 
but 
we 
believe 
that 
the 
changes 
described 
here
are 
straightforward 
and 
obvious, 
given 
the 
existing 
system 
design, 
and
they've 
been 
reviewed 
by 
the 
Clearview 
and 
NWAM 
teams, 
and 
thus 
the
changes 
are 
suitable 
for 
fast-track 
treatment.


1.  
Administration

  
  
All 
of 
the 
administration 
of 
this 
feature 
is 
based 
on 
dladm 
and
  
  
SMF.  
The 
SMF 
portion 
is 
the 
ability 
to 
enable 
and 
disable 
bridge
  
  
instances 
using 
the 
instance 
URIs 
described 
in 
section 
3 
below.

1.1 
New 
dladm 
subcommands

  
  
These 
commands 
are 
patterned 
after 
the 
existing 
aggregation
  
  
commands 
in 
dladm.

  
  
dladm 
create-bridge 
[-t] 
[-R 
<root-dir>] 
[-p 
<priority>]
  
  
  
[-m 
<max-age>] 
[-h 
<hello-time>] 
[-d 
<forward-delay>]
  
  
  
[-f 
<force-protocol>] 
[-l 
<link>]... 
<bridge-name>

  
  
  
This 
command 
creates 
a 
bridge 
instance 
and 
optionally 
assigns
  
  
  
network 
links 
to 
the 
new 
bridge.  
By 
default, 
no 
bridge
  
  
  
instances 
are 
present, 
and 
OpenSolaris 
will 
not 
bridge 
between
  
  
  
network 
links.  
See 
the 
"add-bridge" 
subcommand 
for 
details 
on
  
  
  
link 
assignment.

  
  
  
Bridge 
creation 
and 
link 
assignment 
require 
PRIV_SYS_NET_CONFIG.

  
  
  
In 
order 
to 
bridge 
between 
links, 
you 
must 
create 
at 
least 
one
  
  
  
bridge 
instance.  
Each 
instance 
is 
separate: 
there 
is
  
  
  
intentionally 
no 
forwarding 
connection 
between 
bridges.  
(Note
  
  
  
that 
Crossbow's 
VNICs 
may 
in 
the 
future 
allow 
virtual 
inter-
  
  
  
bridge 
connections.)

  
  
  
The 
<bridge-name> 
provided 
is 
chosen 
by 
the 
administrator 
and
  
  
  
arbitrary, 
but 
must 
be 
a 
legal 
SMF 
service 
instance 
name.  
For
  
  
  
purposes 
of 
documentation, 
this 
is 
a 
URI 
component 
without
  
  
  
escape 
sequences, 
meaning 
that 
the 
following 
characters 
may 
not
  
  
  
be 
present:

    
; 
/ 
? 
: 
@ 
& 
= 
+ 
$ 
, 
% 
< 
> 
# 
"

  
  
  
including 
whitespace 
and 
ASCII 
control 
characters.  
The 
name
  
  
  
"default" 
is 
reserved, 
as 
are 
all 
names 
beginning 
with 
the
  
  
  
string 
"SUNW".  
Names 
with 
trailing 
digits 
are 
not 
permitted, 
in
  
  
  
order 
to 
allow 
for 
creation 
of 
"observability 
devices;" 
see
  
  
  
section 
2 
below.

  
  
  
Because 
of 
the 
use 
of 
the 
observability 
devices, 
the 
names 
of
  
  
  
legal 
bridge 
instances 
are 
further 
constrained 
to 
be 
a 
legal
  
  
  
dlpi(7P) 
name, 
which 
matches:

    
[A-Za-z_][A-Za-z0-9_]*[A-Za-z_]

  
  
  
Options 
are:

  
  
  
-t    
Create 
a 
temporary 
bridge.

    
    
This 
will 
create 
the 
bridge 
object 
on 
the 
running
    
    
system, 
but 
the 
newly 
created 
bridge 
will 
not 
survive
    
    
the 
next 
reboot.

  
  
  
-R 
<root-dir>
    
    
Specify 
an 
alternate 
root 
directory.

    
    
This 
allows 
the 
configuration 
of 
bridge 
instances 
in
    
    
alternate 
roots, 
as 
with 
Live 
Upgrade 
and 
with
    
    
jumpstart 
installs.  
Note 
that 
error 
checking 
for 
link
    
    
type 
isn't 
possible 
when 
administering 
an 
alternate
    
    
root.

  
  
  
-p 
<priority>
    
    
Specify 
the 
Bridge 
Priority.

    
    
This 
sets 
the 
STP 
priority 
value 
for 
determining 
the
    
    
root 
bridge 
node 
in 
the 
network.  
The 
default 
value 
is
    
    
(per 
the 
specification) 
32768, 
and 
legal 
values 
are 
0
    
    
(highest 
priority) 
to 
65535 
(lowest 
priority).

  
  
  
-m 
<max-age>
    
    
Specify 
the 
maximum 
age 
for 
configuration 
information.

    
    
This 
sets 
the 
STP 
Bridge 
Max 
Age 
parameter.
    
    
Information 
older 
than 
this 
(in 
seconds) 
is 
discarded
    
    
by 
all 
bridges 
in 
the 
network 
if 
this 
node 
is 
the 
root
    
    
bridge.  
It 
defaults 
to 
20.0 
seconds.  
Legal 
values
    
    
are 
from 
6.0 
to 
40.0 
seconds.  
(See 
the 
"-d
    
    
<forward-delay>" 
parameter 
for 
additional
    
    
constraints.)

  
  
  
-h 
<hello-time>
    
    
Specify 
the 
Bridge 
Hello 
Time.

    
    
This 
sets 
the 
STP 
Bridge 
Hello 
Time 
parameter.  
If
    
    
this 
node 
is 
the 
root 
node, 
it 
sends 
Configuration
    
    
BPDUs 
at 
this 
interval 
throughout 
the 
network.  
It
    
    
defaults 
to 
2.0 
seconds.  
Legal 
values 
are 
from 
1.0 
to
    
    
10.0 
seconds.  
(See 
the 
"-d 
<forward-delay> 
parameter
    
    
for 
additional 
constraints.)

  
  
  
-d 
<forward-delay>
    
    
Specify 
the 
Bridge 
Forward 
Delay.

    
    
This 
sets 
the 
STP 
Bridge 
Forward 
Delay 
parameter.
    
    
This 
timer 
is 
used 
to 
sequence 
the 
link 
states 
when 
a
    
    
port 
is 
enabled 
anywhere 
in 
the 
network 
if 
this 
node
    
    
is 
the 
root 
bridge.  
It 
defaults 
to 
15.0 
seconds.
    
    
Legal 
values 
are 
from 
4.0 
to 
30.0 
seconds.

    
    
Bridges 
must 
obey 
the 
following 
two 
constraints:

    
    
    
2 
* 
(forward_delay 
- 
1.0) 
>= 
max_age

    
    
    
max_age 
>= 
2 
* 
(hello_time 
+ 
1.0)

    
    
Any 
parameter 
setting 
that 
would 
violate 
those
    
    
constraints 
will 
be 
treated 
as 
an 
error 
and 
cause 
the
    
    
command 
to 
fail 
with 
a 
diagnostic 
message.

  
  
  
-f 
<force-protocol>
    
    
Specify 
the 
forced 
maximum 
supported 
protocol.

    
    
This 
sets 
the 
MSTP 
maximum 
supported 
protocol 
number.
    
    
The 
default 
is 
3.  
The 
current 
implementation 
doesn't
    
    
support 
RSTP 
or 
MSTP, 
so 
this 
currently 
has 
no 
effect.
    
    
However, 
if 
the 
user 
desires 
to 
prevent 
MSTP 
from
    
    
being 
used 
in 
the 
future 
when 
implemented, 
the
    
    
parameter 
may 
be 
set 
to 
0 
(STP 
only) 
or 
2 
(allow
    
    
RSTP).

  
  
  
-l 
<link>    
Add 
a 
link 
to 
the 
newly-created 
bridge.

    
    
This 
is 
equivalent 
to 
creating 
the 
bridge 
and 
then
    
    
adding 
one 
or 
more 
links, 
as 
with 
the 
"add-bridge"
    
    
option 
below, 
except 
that 
if 
any 
of 
the 
links 
cannot
    
    
be 
added, 
then 
the 
entire 
command 
fails, 
and 
the 
new
    
    
bridge 
itself 
isn't 
created.

  
  
dladm 
modify-bridge 
[-t] 
[-R 
<root-dir>] 
[-p 
<priority>]
  
  
  
[-m 
<max-age>] 
[-h 
<hello-time>] 
[-d 
<forward-delay>]
  
  
  
[-f 
<force-protocol>] 
<bridge-name>

  
  
  
This 
subcommand 
modifies 
the 
operational 
parameters 
of 
a 
given
  
  
  
bridge 
instance.  
All 
of 
the 
options 
are 
the 
same 
as 
for 
the
  
  
  
"create-bridge" 
subcommand 
above, 
except 
that 
the 
"-l" 
option 
is
  
  
  
not 
permitted.  
To 
add 
links 
to 
an 
existing 
bridge, 
use 
the
  
  
  
"add-bridge" 
subcommand 
below.

  
  
  
Bridge 
parameter 
modification 
requires 
PRIV_SYS_NET_CONFIG.

  
  
dladm 
delete-bridge 
[-t] 
[-R 
<root-dir>] 
<bridge-name>

  
  
  
This 
subcommand 
deletes 
a 
bridge 
instance.  
Unlike 
the 
bridge
  
  
  
creation 
subcommand, 
which 
can 
add 
links 
while 
creating, 
it 
does
  
  
  
not 
have 
the 
option 
to 
remove 
links 
during 
the 
deletion 
process.
  
  
  
The 
bridge 
being 
deleted 
must 
not 
have 
any 
attached 
links.  
If
  
  
  
it 
does, 
then 
an 
error 
is 
returned 
and 
no 
action 
is 
taken.

  
  
  
Bridge 
deletion 
requires 
PRIV_SYS_NET_CONFIG.

  
  
  
The 
"-t" 
and 
"-R" 
options 
are 
the 
same 
as 
for 
the
  
  
  
"create-bridge" 
subcommand.

  
  
dladm 
add-bridge 
[-t] 
[-R 
<root-dir>] 
-l 
<link> 
[-l 
<link>]...
  
  
  
<bridge-name>

  
  
  
This 
subcommand 
adds 
one 
or 
more 
links 
to 
a 
bridge 
instance.  
If
  
  
  
multiple 
links 
are 
specified, 
and 
adding 
any 
one 
of 
them 
results
  
  
  
in 
an 
error, 
then 
no 
changes 
are 
made 
to 
the 
system 
and 
the
  
  
  
command 
fails.

  
  
  
Link 
addition 
to 
a 
bridge 
requires 
PRIV_SYS_NET_CONFIG.

  
  
  
A 
link 
may 
be 
a 
member 
of 
at 
most 
one 
bridge.  
It's 
an 
error 
to
  
  
  
specify 
that 
a 
link 
belongs 
to 
more 
than 
one 
bridge.  
To 
move 
a
  
  
  
link 
from 
one 
bridge 
instance 
to 
another, 
remove 
it 
from 
the
  
  
  
current 
bridge 
before 
adding 
it 
to 
the 
new 
one.

  
  
  
The 
links 
assigned 
to 
a 
bridge 
must 
not 
themselves 
be 
VLANs 
or
  
  
  
tunnels.  
Only 
links 
that 
would 
be 
acceptable 
as 
part 
of 
an
  
  
  
aggregation 
or 
links 
that 
are 
aggregations 
themselves 
may 
be
  
  
  
assigned 
to 
a 
bridge.  
Other 
link 
types 
will 
result 
in 
error
  
  
  
messages, 
and 
no 
action 
taken.  
(A 
future 
project 
may 
provide
  
  
  
bridging 
over 
tunnels 
using 
GRE, 
and 
over 
PPP 
using 
BCP.  
Those
  
  
  
cases 
are 
not 
part 
of 
this 
project, 
but 
nothing 
this 
project 
is
  
  
  
doing 
will 
preclude 
those 
cases 
from 
the 
future.)

  
  
  
In 
this 
initial 
version, 
the 
links 
must 
also 
be 
Ethernet 
type.
  
  
  
Bridging 
is 
well-defined 
over 
a 
few 
other 
media, 
and 
there 
are
  
  
  
some 
dodgy 
ways 
to 
make 
it 
work 
on 
still 
others, 
but 
those 
cases
  
  
  
are 
subjects 
for 
a 
future 
release.

  
  
  
When 
links 
are 
added 
to 
a 
bridge, 
the 
bridging 
protocol 
in 
use
  
  
  
(STP) 
will 
be 
notified, 
and 
the 
links 
will 
behave 
as 
though 
just
  
  
  
created.  
For 
STP, 
this 
means 
that 
the 
link 
will 
be 
shut 
down
  
  
  
and 
then 
brought 
back 
up 
using 
the 
standard 
protocol.

  
  
  
The 
options 
are 
the 
same 
as 
for 
the 
"create-bridge" 
subcommand.

  
  
dladm 
remove-bridge 
[-t] 
[-R 
<root-dir>] 
-l 
<link> 
[-l 
<link>]...
  
  
  
<bridge-name>

  
  
  
This 
subcommand 
removes 
one 
or 
more 
links 
from 
a 
bridge
  
  
  
instance.  
If 
multiple 
links 
are 
specified, 
and 
removing 
any 
one
  
  
  
of 
them 
would 
result 
in 
an 
error, 
then 
none 
are 
removed 
and 
the
  
  
  
command 
fails.

  
  
  
Link 
removal 
from 
a 
bridge 
requires 
PRIV_SYS_NET_CONFIG.

  
  
  
When 
links 
are 
removed 
from 
a 
bridge, 
the 
bridging 
protocol
  
  
  
(STP) 
is 
notified, 
and 
will 
likely 
recalculate 
a 
new 
network
  
  
  
topology, 
unless 
those 
links 
were 
unused 
due 
to 
loop-pruning
  
  
  
activity 
by 
the 
bridging 
protocol.

  
  
  
The 
options 
are 
the 
same 
as 
for 
the 
"create-bridge" 
subcommand.

  
  
dladm 
show-bridge 
[-p] 
[-s 
[-i 
<interval>]] 
[<bridge-name>]

  
  
  
This 
subcommand 
shows 
the 
running 
status 
of 
bridges.  
When 
given
  
  
  
a 
bridge 
name, 
it 
shows 
the 
status 
of 
that 
one 
bridge.  
If 
no
  
  
  
bridge 
name 
is 
given, 
then 
it 
shows 
summary 
status 
of 
all
  
  
  
bridges 
on 
the 
system.

  
  
  
Note 
the 
lack 
of 
a 
"-R" 
option 
here.  
It 
is 
not 
possible 
to 
list
  
  
  
bridge 
configuration 
information 
in 
an 
alternate 
root, 
in
  
  
  
keeping 
with 
the 
rest 
of 
the 
dladm 
user 
interface.  
The 
reason
  
  
  
for 
this 
restriction 
is 
to 
allow 
the 
data 
to 
be 
represented 
in
  
  
  
SMF, 
where 
"writing" 
to 
an 
alternate 
root 
is 
supported 
by 
way 
of
  
  
  
copying 
appropriate 
commands 
to 
$ROOT/var/svc/profile/upgrade,
  
  
  
but 
"reading" 
is 
not 
feasible 
because 
the 
repository 
on 
the
  
  
  
alternate 
root 
may 
be 
incompatible 
with 
the 
running 
system.

1.2 
New 
dladm 
Link 
Properties

  
  
"stp"

    
This 
is 
a 
boolean 
property.  
It 
defaults 
to 
"true."  
When 
set
    
to 
"false," 
the 
link 
will 
not 
use 
Spanning 
Tree, 
and 
will 
be
    
placed 
into 
forwarding 
mode 
at 
all 
times.  
The 
"false" 
setting
    
is 
appropriate 
for 
point-to-point 
links 
connected 
to 
end
    
nodes.  
Only 
non-VLAN 
type 
links 
have 
this 
property.

  
  
"forward"

    
This 
is 
a 
boolean 
property 
on 
all 
links.  
It 
defaults 
to
    
"true."  
When 
set 
to 
"false," 
the 
VLAN 
associated 
with 
the
    
link 
instance 
will 
not 
forward 
traffic 
through 
the 
bridge.
    
Setting 
the 
property 
to 
"false" 
is 
equivalent 
to 
removing 
the
    
VLAN 
from 
the 
"allowed 
set" 
for 
a 
traditional 
bridge.

  
  
"default-tag"

    
This 
is 
a 
numeric 
property 
with 
range 
0 
to 
4094.  
It 
defaults
    
to 
1.  
It 
defines 
the 
default 
VLAN 
ID 
that's 
assumed 
for
    
untagged 
packets 
sent 
to 
and 
received 
from 
this 
link.  
Only
    
non-VLAN 
type 
links 
have 
this 
property.

  
  
"stp-priority"

    
This 
is 
a 
numeric 
property 
with 
range 
0 
to 
255.  
It 
defaults
    
to 
128.  
It 
corresponds 
to 
the 
STP 
Port 
Priority 
value, 
which
    
is 
used 
to 
determine 
the 
preferred 
root 
port 
on 
a 
bridge 
by
    
prepending 
to 
the 
port 
identifier.  
Lower 
numerical 
values 
are
    
higher 
priority.

  
  
"stp-cost"

    
This 
is 
a 
numeric 
property 
with 
range 
1 
to 
65535; 
zero 
is 
not
    
allowed.  
It 
represents 
the 
cost 
for 
using 
the 
link, 
and
    
defaults 
to 
(per 
the 
standard) 
100 
for 
10Mbps, 
19 
for 
100Mbps,
    
4 
for 
1Gbps, 
and 
2 
for 
10Gbps.

  
  
"bridge-port"

    
This 
is 
a 
read-only 
numeric 
property.  
It 
shows 
the 
port
    
number 
for 
the 
link 
as 
seen 
by 
the 
bridge, 
and 
is 
used 
in
    
Spanning 
Tree 
messages 
and 
network 
management.

1.3 
New 
Kstats

  
  
Each 
bridge 
instance 
will 
have 
a 
set 
of 
statistics, 
named
  
  
"bridge:<index>:<bridge-name>:<statistic>", 
where:

    
<index>
    
    
Arbitrary 
instance 
number 
assigned 
by 
the 
kernel 
and
    
    
not 
necessarily 
retained 
across 
reboot.

    
<bridge-name>
    
    
Administrator-specified 
bridge 
name.

    
<statistic>
    
    
Name 
of 
statistic; 
at 
least 
the 
following:

    
    
learn_source    
Number 
of 
sources 
learned
    
    
learn_expire    
Number 
of 
learnt 
entries 
expired
    
    
learn_size    
Current 
count 
of 
learnt 
entries
    
    
forward_direct    
Directly 
forwarded 
packet 
count
    
    
forward_unknown    
Forwarded 
with 
unknown 
destination
    
    
forward_mbcast    
Forwarded 
multicast/broadcast

  
  
Each 
link 
instance 
will 
also 
have 
new 
kstats, 
where 
the
  
  
<statistic> 
names 
will 
be:

    
bridge_sent    
Packets 
forwarded 
to 
the 
link 
by 
bridging
    
bridge_rcvd    
Packets 
received 
from 
the 
link 
(and 
forwarded
    
    
    
elsewhere) 
by 
bridging

  
  
All 
of 
these 
statistics 
are 
considered 
Volatile 
for 
now.  
The
  
  
existence 
of 
the 
statistics 
will 
be 
documented 
for 
users, 
but 
with
  
  
warnings 
that 
the 
names 
and 
definitions 
of 
the 
statistics 
may
  
  
change 
incompatibly.  
A 
future 
case 
for 
the 
overall 
RBridges
  
  
project 
will 
elevate 
these 
in 
stability.


2.  
Packet 
Observability

  
  
Each 
bridge 
instance 
will 
be 
assigned 
an 
"observability 
device,"
  
  
in 
a 
manner 
similar 
to 
the 
DLPI 
nodes 
created 
for 
"Clearview: 
IP
  
  
Observability 
Devices" 
(PSARC 
2006/475).  
These 
nodes 
will 
appear
  
  
under 
the 
/dev/bridge/ 
directory, 
named 
by 
the 
bridge 
name 
plus 
a
  
  
trailing 
"0".

  
  
The 
observability 
node 
is 
intended 
for 
use 
with 
snoop 
and
  
  
wireshark.  
It 
behaves 
as 
a 
standard 
Ethernet 
interface, 
but 
does
  
  
not 
permit 
the 
transmission 
of 
packets.  
All 
transmitted 
packets
  
  
are 
silently 
dropped.

  
  
The 
user 
of 
this 
node 
will 
get 
a 
single 
unmodified 
copy 
of 
every
  
  
packet 
handled 
by 
the 
bridge, 
similar 
to 
a 
"monitoring" 
port 
on 
a
  
  
traditional 
bridge, 
and 
subject 
to 
the 
usual 
DLPI 
"promiscuous
  
  
mode" 
rules.  
The 
user 
may 
also 
filter 
on 
VLAN 
ID 
by 
using 
the
  
  
VLAN 
PPA 
hack 
mechanism: 
"/dev/bridge/my-bridge1000" 
selects 
VLAN
  
  
ID 
1 
on 
bridge 
the 
instance 
named 
"my-bridge".

  
  
The 
observability 
node 
also 
forms 
a 
Project 
Private 
control 
node
  
  
for 
the 
kernel, 
allowing 
ioctls 
to 
a 
specific 
bridge 
instance, 
and
  
  
will 
be 
used 
by 
the 
STP 
daemon 
and 
other 
(future) 
bridging
  
  
protocols.

  
  
The 
dlpi_open(3DLPI) 
interface 
will 
be 
enhanced 
with 
a 
Committed
  
  
DLPI_BRIDGE 
flag 
to 
allow 
applications 
to 
locate 
the 
observability
  
  
nodes 
by 
name.


3.  
STP 
Daemon

  
  
Each 
bridge 
(created 
via 
"dladm 
create-bridge") 
is 
represented 
as
  
  
an 
identically-named 
SMF 
instance 
of 
svc:/network/bridge.  
Each
  
  
instance 
runs 
a 
copy 
of 
/usr/lib/bridged, 
which 
implements 
the
  
  
Spanning 
Tree 
Protocol 
(STP).  
For 
example, 
if 
the 
user 
runs:

    
# 
dladm 
create-bridge 
my-bridge

  
  
The 
system 
will 
have 
an 
SMF 
service 
named:

    
svc:/network/bridge:my-bridge

  
  
and 
(per 
section 
2 
above) 
an 
observability 
node 
named:

    
/dev/bridge/my-bridge0

  
  
By 
default, 
all 
ports 
run 
standard 
STP.  
This 
is 
done 
for 
safety
  
  
reasons: 
a 
bridge 
that 
does 
not 
run 
some 
form 
of 
bridging 
protocol
  
  
(such 
as 
STP) 
can 
form 
long-lasting 
forwarding 
loops 
in 
the
  
  
network.  
Because 
Ethernet 
has 
no 
hop-count 
or 
TTL 
on 
packets, 
any
  
  
such 
loops 
are 
fatal 
to 
the 
network.

  
  
When 
the 
adminstrator 
knows 
that 
a 
particular 
port 
is 
not
  
  
connected 
to 
another 
bridge 
(for 
example, 
a 
direct 
point-to-point
  
  
connection 
to 
a 
host 
system), 
STP 
can 
be 
disabled 
administratively
  
  
for 
that 
port.  
Even 
if 
all 
ports 
on 
a 
bridge 
have 
STP 
disabled,
  
  
the 
STP 
daemon 
still 
runs; 
this 
is 
in 
case 
new 
ports 
are 
added,
  
  
and 
because 
it 
is 
responsible 
for 
enabling 
and 
disabling
  
  
forwarding 
on 
the 
ports.

  
  
If 
the 
SMF 
service 
instance 
for 
a 
bridge 
is 
disabled, 
then 
bridge
  
  
forwarding 
stops 
on 
those 
ports 
as 
the 
STP 
daemon 
is 
stopped.  
If
  
  
the 
instance 
is 
restarted, 
STP 
starts 
from 
its 
initial 
state.

  
  
The 
bridge 
daemon 
runs 
as 
UID/GID 
"daemon" 
with
  
  
PRIV_SYS_NET_CONFIG 
in 
order 
to 
access 
the 
raw 
network 
devices,
  
  
but 
with 
most 
other 
basic 
privileges 
(e.g., 
PRIV_PROC_FORK 
and
  
  
PRIV_PROC_EXEC) 
removed.


3.  
VLANs

  
  
In 
general, 
administrators 
will 
want 
to 
have 
the 
VLANs 
they
  
  
configure 
on 
the 
system 
to 
be 
forwarded 
among 
all 
the 
ports 
on 
a
  
  
bridge 
instance, 
so 
this 
will 
be 
the 
default 
for 
VLANs.  
When 
the
  
  
administrator 
invokes 
Clearview's 
"dladm 
create-vlan", 
and 
the
  
  
underlying 
link 
is 
part 
of 
a 
bridge, 
that 
command 
will 
also 
enable
  
  
forwarding 
of 
the 
specified 
VLAN 
on 
that 
bridge 
link.

  
  
If 
an 
administrator 
wants 
to 
configure 
a 
VLAN 
on 
a 
link 
but 
not
  
  
allow 
forwarding 
to 
or 
from 
other 
links 
on 
the 
bridge, 
then 
he
  
  
must 
take 
specific 
action 
to 
do 
so, 
by 
disabling 
forwarding 
with
  
  
"set-linkprop".

  
  
Clearview 
UV 
provides 
two 
mechanisms 
for 
the 
creation 
of 
VLANs.
  
  
The 
primary 
means 
of 
configuration 
is 
the 
new 
"dladm 
create-vlan"
  
  
subcommand, 
which 
automatically 
enables 
the 
VLAN 
for 
bridging 
as
  
  
described 
above, 
if 
the 
underlying 
link 
is 
configured 
as 
part 
of 
a
  
  
bridge.

  
  
The 
second 
mechanism 
is 
a 
legacy 
feature 
called 
the 
"PPA 
hack."
  
  
This 
allows 
a 
user 
to 
create 
a 
VLAN 
simply 
by 
opening 
a 
DLPI
  
  
provider 
and 
specifying 
a 
VLAN 
ID 
number 
as 
part 
of 
the 
PPA.  
In
  
  
this 
case, 
the 
user 
may 
be 
doing 
nothing 
other 
than 
snooping 
on
  
  
that 
VLAN, 
so 
adding 
the 
VLAN 
to 
the 
allowed 
set 
automatically 
is
  
  
likely 
not 
the 
right 
answer.  
Thus, 
we 
will 
default 
forwarding 
to
  
  
"off" 
for 
PPA-hack 
VLANs.  
Administrators 
with 
legacy 
PPA 
hack
  
  
VLANs 
will 
need 
to 
reconfigure 
to 
use 
the 
new 
Clearview 
VLANs 
to
  
  
take 
full 
advantage 
of 
bridging, 
and 
this 
will 
be 
included 
in 
the
  
  
documentation.

  
  
In 
STP, 
VLANs 
are 
ignored.  
The 
bridging 
protocol 
computes 
just
  
  
one 
loop-free 
topology 
and 
uses 
that.  
Administrators 
are 
required
  
  
to 
configure 
any 
"duplicate" 
links 
such 
that 
when 
they're
  
  
automatically 
disabled 
by 
STP, 
the 
configured 
VLANs 
are 
not
  
  
disconnected.  
MSTP 
is 
somewhat 
similar, 
but 
allows 
administrators
  
  
to 
assign 
each 
VLAN 
to 
a 
small 
number 
of 
distinct 
spanning 
tree
  
  
"instances," 
and 
allows 
instances 
within 
an 
identically-configured
  
  
"region" 
to 
have 
distinct 
topologies.  
In 
terms 
of 
this 
project,
  
  
additional 
bridge 
and 
link 
properties 
would 
be 
required 
to 
enable
  
  
MSTP 
operation.


4.  
SMF 
Properties

  
  
These 
parameters 
are 
all 
Project 
Private.  
They 
will 
not 
be
  
  
documented, 
and 
the 
documented 
administrative 
interface 
will 
be
  
  
the 
dladm 
command.

4.1 
STP 
SMF

  
  
Property 
Name    
    
Type    
    
Default
  
  
--------------    
    
----    
    
-------
  
  
config/priority    
    
ushort_t    
32768
  
  
config/max-age    
    
ushort_t    
5120    
(20 
seconds)
  
  
config/hello-time    
    
ushort_t    
512    
(2 
seconds)
  
  
config/forward-delay    
ushort_t    
3840    
(15 
seconds)
  
  
config/force-protocol    
int    
    
3

  
  
All 
of 
these 
properties 
(and 
their 
default 
values 
and
  
  
granularities) 
are 
defined 
by 
the 
STP 
and 
related 
standards.

  
  
The 
"force-protocol" 
parameter 
is 
specified 
to 
allow 
for 
an
  
  
upgrade 
path.  
Users 
who 
do 
not 
want 
to 
see 
the 
use 
of 
MSTP 
when
  
  
it 
is 
implemented 
can 
set 
this 
parameter 
to 
0 
or 
2 
(as 
specified
  
  
in 
IEEE 
802.1Q-2004) 
to 
select 
STP 
or 
RSTP 
as 
the 
maximum 
allowed
  
  
protocol.  
In 
this 
project, 
the 
parameter 
will 
have 
no 
effect, 
as
  
  
only 
STP 
is 
implemented.

4.2 
Datalink 
SMF

  
  
Property 
Name    
    
Type    
    
Default
  
  
--------------    
    
----    
    
-------
  
  
config/stp    
    
    
boolean    
    
true
  
  
config/forward    
    
boolean    
    
true
  
  
config/bridge    
    
string    
    
""
  
  
config/default-tag    
    
ushort_t    
1

  
  
On 
a 
Nemo 
device, 
legacy 
device, 
or 
aggregation, 
the 
link
  
  
parameters 
are 
used 
as 
above.  
The 
"default-tag" 
parameter 
may 
be
  
  
set 
to 
0 
to 
disable 
the 
forwarding 
of 
untagged 
packets 
to 
and 
from
  
  
the 
port.

  
  
On 
a 
VLAN, 
"stp" 
and 
"default-tag" 
are 
ignored.  
The 
"forward"
  
  
flag 
enables 
forwarding 
for 
that 
VLAN, 
which 
is 
equivalent 
to
  
  
putting 
the 
VLAN 
into 
the 
"allowed 
set" 
for 
the 
bridge 
port.
  
  
Setting 
it 
to 
"false" 
causes 
the 
VLAN 
to 
be 
disallowed, 
which
  
  
means 
that 
VLAN-based 
I/O 
to 
the 
underlying 
link 
still 
operates,
  
  
but 
no 
bridge-based 
forwarding 
is 
done.  
The 
"bridge" 
parameter 
is
  
  
reserved 
for 
use 
with 
MSTP, 
where 
it 
will 
select 
an 
instance.


5.  
Alternatives

5.1 
Using 
A 
Separate 
Command

  
  
An 
alternative 
command 
set 
design 
would 
be 
to 
create 
a 
new 
bridge
  
  
control 
command 
(bridgeadm), 
rather 
than 
using 
dladm.

  
  
The 
main 
problem 
with 
this 
separation 
is 
that 
the 
configuration 
of
  
  
the 
bridge 
would 
end 
up 
being 
split 
between 
two 
different
  
  
utilities 
in 
a 
somewhat 
incoherent 
manner.  
Why 
would 
IEEE 
802
  
  
aggregations 
be 
part 
of 
dladm 
but 
IEEE 
802 
bridges 
be 
configured
  
  
elsewhere?

  
  
Parts 
of 
the 
configuration 
of 
a 
bridge 
(such 
as 
the 
set 
of 
allowed
  
  
VLANs 
and 
the 
default 
VLAN 
tag 
for 
a 
given 
link) 
are 
naturally
  
  
part 
of 
the 
link 
configuration, 
and 
not 
a 
common 
property 
of 
the
  
  
bridge.  
The 
creation 
of 
VLANs 
(logically 
located 
"above" 
links
  
  
and 
bridges) 
and 
regular 
Ethernet 
links 
(logically 
located 
"below"
  
  
VLANs 
and 
bridges) 
via 
dladm 
while 
bridging 
itself 
is 
in 
bridgeadm
  
  
seems 
like 
a 
very 
strange 
result.

  
  
We 
could 
create 
a 
separate 
bridgeadm, 
but 
then 
we'd 
likely 
have 
to
  
  
deal 
with 
the 
VLAN 
issues 
some 
other 
way.  
Most 
likely, 
we 
would
  
  
end 
up 
with 
either 
duplicate 
configuration 
in 
bridgeadm 
or 
the
  
  
bulk 
of 
bridge 
configuration 
actually 
going 
on 
in 
dladm 
per-link
  
  
properties, 
and 
only 
bridge 
create/destroy 
done 
via 
bridgeadm.

  
  
In 
other 
words, 
there 
are 
several 
IEEE-specified 
parameters 
for
  
  
bridges, 
but 
they're 
rarely 
of 
much 
interest, 
so 
that 
proposed
  
  
utility 
wouldn't 
do 
very 
much.  
The 
main 
thing 
users 
need 
to
  
  
manipulate 
for 
bridges 
are 
the 
VLANs, 
and 
we 
need 
to 
figure 
out
  
  
how 
to 
represent 
that 
manipulation.  
We 
choose 
to 
equate 
dladm-
  
  
created-VLAN 
with 
bridge-allowed-VLAN 
because 
it 
seems 
to 
produce
  
  
the 
most 
natural 
results: 
there's 
only 
one 
way 
to 
"create" 
or
  
  
"destroy" 
a 
VLAN 
in 
the 
system.

  
  
The 
alternative 
is 
to 
break 
those 
apart, 
and 
allow 
users 
to 
create
  
  
VLANs 
for 
potential 
use 
with 
IP 
via 
dladm, 
and 
separately 
assign
  
  
VLANs 
to 
bridge 
ports 
via 
bridgeadm, 
but 
that 
runs 
the 
very 
likely
  
  
risk 
of 
misconfiguration: 
either 
forgetting 
to 
enable 
a 
bridge
  
  
link 
for 
a 
VLAN 
while 
having 
IP 
plumbed 
atop, 
or 
thinking 
that
  
  
destroying 
the 
VLAN 
removes 
it 
from 
the 
bridge.  
Since 
neither
  
  
scenario 
seems 
to 
be 
particularly 
useful, 
allowing 
for 
them
  
  
doesn't 
seem 
like 
a 
good 
goal.

  
  
Or, 
for 
a 
really 
short 
answer: 
dladm 
is 
the 
location 
of 
all 
things
  
  
datalinkish, 
and 
bridging 
is 
(like 
VLANs 
and 
aggregations) 
a
  
  
datalink 
function.

5.2 
Link 
Configuration 
Storage

  
  
Alternative 
designs 
for 
the 
configuration 
information 
include
  
  
having 
the 
set 
of 
links 
for 
a 
bridge 
listed 
as 
part 
of 
the 
bridge
  
  
configuration, 
and 
using 
non-SMF 
files 
for 
storing 
configuration.

  
  
The 
former 
approach 
would 
work, 
and 
would 
have 
the 
advantage 
that
  
  
during 
start-up 
of 
the 
STP 
daemon 
it 
would 
be 
easy 
to 
find 
the
  
  
list 
of 
links 
configured 
for 
that 
instance.  
That's 
a 
benefit 
over
  
  
the 
proposed 
design 
in 
that 
we 
will 
need 
to 
iterate 
over 
all 
links
  
  
to 
get 
the 
list 
needed 
for 
a 
single 
instance.  
However, 
there 
are
  
  
two 
reasons 
this 
approach 
wasn't 
chosen:

    
a. 
A 
link 
may 
be 
a 
member 
of 
at 
most 
one 
bridge.  
This
    
  
 
semantic 
is 
easy 
to 
enforce 
with 
a 
link 
property, 
as
    
  
 
there's 
just 
one 
instance 
of 
the 
property, 
but 
is 
hard 
to
    
  
 
enforce 
across 
multiple 
bridges.  
We 
end 
up 
needing 
to 
scan
    
  
 
all 
bridge 
instances, 
and 
configuration 
transactions 
become
    
  
 
more 
complex 
because 
two 
objects 
need 
to 
be 
changed 
at 
one
    
  
 
time.

    
b. 
We 
want 
to 
have 
all 
configuration 
parameters 
for 
a 
link 
to
    
  
 
be 
stored 
with 
the 
link 
itself.  
Having 
parameters 
stored
    
  
 
elsewhere 
in 
the 
system 
means 
that 
utilities 
that
    
  
 
manipulate 
links 
or 
just 
display 
system 
configuration 
may
    
  
 
end 
up 
needing 
to 
scan 
through 
these 
other 
locations 
in
    
  
 
order 
to 
make 
coherent 
system 
changes.  
(For 
this 
project,
    
  
 
we 
would 
be 
forced 
to 
change 
the 
existing 
Clearview 
"dladm
    
  
 
delete-link" 
functionality 
so 
that 
it 
scanned 
the 
bridge
    
  
 
instances 
and 
removed 
any 
links 
found 
there.  
Storing 
the
    
  
 
data 
with 
the 
link 
instance 
removes 
that 
requirement.)

  
  
Using 
non-SMF 
files 
would 
also 
work, 
and 
we 
could 
make 
use 
of 
the
  
  
Clearview 
UV 
"link 
IDs" 
to 
avoid 
problems 
inherent 
with 
link
  
  
renaming.  
However, 
longer 
term, 
the 
Clearview 
and 
NWAM 
teams 
are
  
  
refactoring 
link 
configuration 
into 
SMF.  
Having 
native 
bridging
  
  
designed 
for 
OpenSolaris 
but 
not 
actually 
integrated 
with 
its 
core
  
  
administrative 
mechanisms 
seems 
like 
a 
poor 
recipe 
for 
the 
future.


6.  
Interface 
Summary

  
  
Interface    
    
Stability    
    
Comments
  
  
---------    
    
---------    
    
--------
  
  
dladm 
*-bridge    
Committed
  
  
link 
properties    
Committed
  
  
kstats    
    
Volatile    
    
Should 
be 
raised 
later
  
  
/dev/bridge/    
Committed    
    
Observability 
node
  
  
control 
ioctls    
Project 
Private
  
  
/usr/lib/bridged    
Project 
Private
  
  
/network/bridge    
Committed    
    
SMF 
URI
  
  
config/*    
    
Project 
Private    
    
SMF 
properties
  
  
bridge 
module    
Project 
Private    
    
Kernel 
bridging 
module
  
  
DLPI_BRIDGE    
    
Committed    
    
dlpi_open(3DLPI)
_______________________________________________
opensolaris-arc 
mailing 
list
opensolaris-arc@opensolaris.org





      ____________________________________________________________________________________
Never miss a thing.  Make Yahoo your home page. 
http://www.yahoo.com/r/hs

From Nicolas.Williams@sun.com Mon Jan 28 14:21:40 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0SMLdOp024603
	for <psarc-ext@sac.sfbay.sun.com>; Mon, 28 Jan 2008 14:21:40 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m0SMLd7Q032581;
	Mon, 28 Jan 2008 15:21:39 -0700 (MST)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVD00603KS24500@nwk-avmta-2.sfbay.sun.com>; Mon,
 28 Jan 2008 14:21:38 -0800 (PST)
Received: from binky.Central.Sun.COM ([129.153.128.104])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVD008FOKS1MAD0@nwk-avmta-2.sfbay.sun.com>; Mon,
 28 Jan 2008 14:21:37 -0800 (PST)
Received: from binky.Central.Sun.COM (localhost [127.0.0.1])
	by binky.Central.Sun.COM (8.14.1+Sun/8.14.1) with ESMTP id m0SMLbsX014874;
 Mon, 28 Jan 2008 16:21:37 -0600 (CST)
Received: (from nw141292@localhost)
	by binky.Central.Sun.COM (8.14.1+Sun/8.14.1/Submit) id m0SMLbfQ014873; Mon,
 28 Jan 2008 16:21:37 -0600 (CST)
Date: Mon, 28 Jan 2008 16:21:37 -0600
From: Nicolas Williams <Nicolas.Williams@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <18334.20484.557823.2174@gargle.gargle.HOWL>
To: James Carlson <James.D.Carlson@sun.com>
Cc: psarc-ext@sun.com
Mail-followup-to: James Carlson <James.D.Carlson@Sun.COM>, psarc-ext@sun.com
Message-id: <20080128222136.GP12865@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
Content-disposition: inline
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
X-Authentication-warning: binky.Central.Sun.COM: nw141292 set sender to
 Nicolas.Williams@sun.com using -f
User-Agent: Mutt/1.5.7i
Status: RO
Content-Length: 571

On Mon, Jan 28, 2008 at 04:58:28PM -0500, James Carlson wrote:
>     dladm create-bridge [-t] [-R <root-dir>] [-p <priority>]
>       [-m <max-age>] [-h <hello-time>] [-d <forward-delay>]
>       [-f <force-protocol>] [-l <link>]... <bridge-name>

Will the future RBRIDGES project re-use the *-bridge commands or will it
introduce new *-rbridge commands?  If the former, shouldn't all those
STP-specific getops options be treated as properties instead because
they may not have equivalents in RBRIDGE?  (Or for that matter, RSTP and
MSTP, but I know nothing about them.)

From carlsonj@phorcys.east.sun.com Tue Jan 29 04:30:26 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0TCUPEM013334
	for <psarc-ext@sac.sfbay.Sun.COM>; Tue, 29 Jan 2008 04:30:26 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id m0TCUNlF005132
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Tue, 29 Jan 2008 20:30:24 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVE0021PO2N2O00@nwk-avmta-1.sfbay.Sun.COM> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Tue, 29 Jan 2008 04:30:23 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVE00EN4O2KX2D0@nwk-avmta-1.sfbay.Sun.COM> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Tue,
 29 Jan 2008 04:30:20 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m0TCUJrI007474; Tue,
 29 Jan 2008 07:30:19 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m0TCUJl7007471; Tue,
 29 Jan 2008 07:30:19 -0500 (EST)
Date: Tue, 29 Jan 2008 07:30:19 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <401216.62773.qm@web30801.mail.mud.yahoo.com>
To: Octave Orgeron <unixconsole@yahoo.com>
Cc: psarc-ext@sun.com
Message-id: <18335.7259.731379.532628@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <401216.62773.qm@web30801.mail.mud.yahoo.com>
Status: RO
Content-Length: 1934

Octave Orgeron writes:
> Hi James,
> 
> It's good to see this come up for integration.

Just to set expectations: we're actually nowhere _near_ integration.
That's months away.  We're still in an early stage of design and
development, and we're beginning to nail down some details.  If you're
interested in the project in more depth:

  http://www.opensolaris.org/os/project/rbridges/

> I was wondering how it'll impact virtualization (containers, xen, and ldoms)? Will further work be required or will it work transparently with things like VSW's in LDoms?

In general, it follows the same model as 802.3 aggregation.  Any place
that can be configured, bridging can be configured, and where it
can't, you won't have bridges.

For Zones, bridges won't be possible.  The default shared-stack type
of zone does not have access to datalink layer interfaces at all, so
bridging wouldn't apply.  For the exclusive-stack type of zone, the
current assignment logic places only VLAN objects inside the zone, and
not the basic links, so you'll have no way to construct a bridge.

(If you could do that, the semantics would be strange.  You'd end up
with one zone connecting interfaces together behind the collective
backs of the other zones, and little in the way of observability.  I
suppose it's possible, but the system [particularly the Nemo link
layer] isn't currently designed to support it.)

For Xen and LDOMs, I'd expect bridging to work normally, in as much as
those solutions give the guest operating system access to a MAC-layer
device of some sort.  Obviously, there are some permutations in here
that we'll need to include in our testing, but I don't see any
architectural issues that would preclude it.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From carlsonj@phorcys.east.sun.com Tue Jan 29 05:19:36 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0TDJZZW013927
	for <psarc-ext@sac.sfbay.Sun.COM>; Tue, 29 Jan 2008 05:19:36 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id m0TDJUEL022551;
	Tue, 29 Jan 2008 21:19:33 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVE00601QCIV700@nwk-avmta-1.sfbay.Sun.COM>; Tue,
 29 Jan 2008 05:19:30 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVE0069QQCIL710@nwk-avmta-1.sfbay.Sun.COM>; Tue,
 29 Jan 2008 05:19:30 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m0TDJT5Y007554; Tue,
 29 Jan 2008 08:19:29 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m0TDJTJf007551; Tue,
 29 Jan 2008 08:19:29 -0500 (EST)
Date: Tue, 29 Jan 2008 08:19:29 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <20080128222136.GP12865@Sun.COM>
To: Nicolas Williams <Nicolas.Williams@sun.com>
Cc: psarc-ext@sun.com
Message-id: <18335.10209.804433.285874@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <20080128222136.GP12865@Sun.COM>
Status: RO
Content-Length: 2189

Nicolas Williams writes:
> On Mon, Jan 28, 2008 at 04:58:28PM -0500, James Carlson wrote:
> >     dladm create-bridge [-t] [-R <root-dir>] [-p <priority>]
> >       [-m <max-age>] [-h <hello-time>] [-d <forward-delay>]
> >       [-f <force-protocol>] [-l <link>]... <bridge-name>
> 
> Will the future RBRIDGES project re-use the *-bridge commands or will it
> introduce new *-rbridge commands?

The former.

>  If the former, shouldn't all those
> STP-specific getops options be treated as properties instead because
> they may not have equivalents in RBRIDGE?

That should not be necessary, but it's a good question.  RBridges also
need to be able to act as traditional bridges in order to allow for
mixed-mode networks, and because some of the detailed operation of
TRILL requires STP data.  (In particular, you need to know the root
bridge node ID in order to detect whether there are RBridge links
attached to the same externally-bridged network but where filtering or
VLAN configuration prevents direct communication between the RBridges
themselves.)

We may need to add one or two new parameters here (not quite scoped
yet), but the STP ones should stay.

>  (Or for that matter, RSTP and
> MSTP, but I know nothing about them.)

RSTP is mostly a minor refinement on STP.  MSTP presents a more
substantial change: it allows the administrator to create a handful of
STP instances, and then manually assign VLANs to these instances.  The
idea is to allow for a few different possible topologies within the
network that somehow form the basis for selected sets of VLANs.

It looks pretty hard to administer to me, and substantially less
capable and flexible than RBridges.  I'm not entirely sure why you'd
want it, except perhaps to be buzzword compliant.  But if someone does
want it, then we may eventually need to do it.  The effect on the
configuration would be to need some way of assigning VLANs to STP
instance, which is what's outlined in this document.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From sowmini.varadhan@sun.com Tue Jan 29 06:00:37 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0TE0aei014578
	for <psarc-ext@sac.sfbay.sun.com>; Tue, 29 Jan 2008 06:00:36 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id m0TE0UJd004211
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Tue, 29 Jan 2008 14:00:35 GMT
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVE0023LS8YB000@brm-avmta-1.central.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Tue, 29 Jan 2008 07:00:34 -0700 (MST)
Received: from dm-east-02.east.sun.com ([129.148.13.5])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVE00DI9S8XO8A0@brm-avmta-1.central.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Tue,
 29 Jan 2008 07:00:33 -0700 (MST)
Received: from quasimodo.East.Sun.COM (quasimodo.East.Sun.COM [129.148.174.94])
	by dm-east-02.east.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2)
 with ESMTP id m0TE0WKT055565; Tue, 29 Jan 2008 09:00:32 -0500 (EST)
Received: from quasimodo.East.Sun.COM (localhost [127.0.0.1])
	by quasimodo.East.Sun.COM (8.14.1+Sun/8.14.1) with ESMTP id m0TDqVGe004510;
 Tue, 29 Jan 2008 08:52:31 -0500 (EST)
Received: (from sowmini@localhost)	by quasimodo.East.Sun.COM
 (8.14.1+Sun/8.14.1/Submit) id m0TDqVvt004509; Tue,
 29 Jan 2008 08:52:31 -0500 (EST)
Date: Tue, 29 Jan 2008 08:52:31 -0500
From: sowmini.varadhan@sun.com
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <18335.10209.804433.285874@gargle.gargle.HOWL>
To: James Carlson <James.D.Carlson@sun.com>
Cc: psarc-ext@sun.com
Message-id: <20080129135231.GM2418@quasimodo.East.Sun.COM>
MIME-version: 1.0
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
Content-disposition: inline
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <20080128222136.GP12865@Sun.COM> <18335.10209.804433.285874@gargle.gargle.HOWL>
X-Authentication-warning: quasimodo.East.Sun.COM: sowmini set sender to
 sowmini.varadhan@sun.com using -f
User-Agent: Mutt/1.5.16 (2007-06-11)
Status: RO
Content-Length: 233


Jim,

a nit:

> dladm show-bridge [-p] [-s [-i <interval>]] [<bridge-name>]

should support the -o option, like the other show commands, i.e.,

  dladm show-bridge [-p] [-o field,..] [-s [-i <interval>]] [<bridge-name>]

--Sowmini


From carlsonj@phorcys.east.sun.com Tue Jan 29 06:40:48 2008
Received: from sunmail3mpk.sfbay.sun.com (sunmail3mpk [129.146.11.52])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0TEelni015116
	for <psarc-ext@sac.sfbay.sun.com>; Tue, 29 Jan 2008 06:40:47 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail3mpk.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m0TEekfh017805;
	Tue, 29 Jan 2008 06:40:47 -0800 (PST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVE00503U3VIF00@brm-avmta-1.central.sun.com>; Tue,
 29 Jan 2008 07:40:43 -0700 (MST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVE00D77U3VO8E0@brm-avmta-1.central.sun.com>; Tue,
 29 Jan 2008 07:40:43 -0700 (MST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m0TEegVl007966; Tue,
 29 Jan 2008 09:40:42 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m0TEegsZ007963; Tue,
 29 Jan 2008 09:40:42 -0500 (EST)
Date: Tue, 29 Jan 2008 09:40:42 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <20080129135231.GM2418@quasimodo.East.Sun.COM>
To: Sowmini.Varadhan@sun.com
Cc: PSARC-ext@sun.com
Message-id: <18335.15082.901089.140848@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <20080128222136.GP12865@Sun.COM>
 <18335.10209.804433.285874@gargle.gargle.HOWL>
 <20080129135231.GM2418@quasimodo.East.Sun.COM>
Status: RO
Content-Length: 646

Sowmini.Varadhan@Sun.COM writes:
> a nit:
> 
> > dladm show-bridge [-p] [-s [-i <interval>]] [<bridge-name>]
> 
> should support the -o option, like the other show commands, i.e.,
> 
>   dladm show-bridge [-p] [-o field,..] [-s [-i <interval>]] [<bridge-name>]

Good point.  I'd modeled that one after the existing show-aggr, but I
should have picked up the clearview/brussels change.  Consider the
proposal updated.  ;-}

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From carlsonj@phorcys.east.sun.com Tue Jan 29 08:05:42 2008
Received: from sunmail2sca.sfbay.sun.com (sunmail2sca [129.145.155.234])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0TG5gSb017165
	for <psarc-ext@sac.sfbay.sun.com>; Tue, 29 Jan 2008 08:05:42 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail2sca.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m0TG5fE4001799;
	Tue, 29 Jan 2008 08:05:41 -0800 (PST)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVE00M0RY1H2800@nwk-avmta-1.sfbay.Sun.COM>; Tue,
 29 Jan 2008 08:05:41 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVE00G1BY1FH780@nwk-avmta-1.sfbay.Sun.COM>; Tue,
 29 Jan 2008 08:05:39 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m0TG5cDT008648; Tue,
 29 Jan 2008 11:05:38 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m0TG5cif008645; Tue,
 29 Jan 2008 11:05:38 -0500 (EST)
Date: Tue, 29 Jan 2008 11:05:38 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <18335.15082.901089.140848@gargle.gargle.HOWL>
To: Sowmini.Varadhan@sun.com, PSARC-ext@sun.com
Message-id: <18335.20178.808784.163318@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <20080128222136.GP12865@Sun.COM>
 <18335.10209.804433.285874@gargle.gargle.HOWL>
 <20080129135231.GM2418@quasimodo.East.Sun.COM>
 <18335.15082.901089.140848@gargle.gargle.HOWL>
Status: RO
Content-Length: 4704

James Carlson writes:
> Sowmini.Varadhan@Sun.COM writes:
> > a nit:
> > 
> > > dladm show-bridge [-p] [-s [-i <interval>]] [<bridge-name>]
> > 
> > should support the -o option, like the other show commands, i.e.,
> > 
> >   dladm show-bridge [-p] [-o field,..] [-s [-i <interval>]] [<bridge-name>]
> 
> Good point.  I'd modeled that one after the existing show-aggr, but I
> should have picked up the clearview/brussels change.  Consider the
> proposal updated.  ;-}

Your comment made me go back and update the spec in detail to show the
parameters that can be shown, and to add support for showing the
running bridging-related port status described in 802.1D, as well as
private interfaces used.  The diffs against the original spec are
below, and the new spec is in the case directory as "spec.txt".

diff -r 5243e2e65184 bridging-spec.txt
--- a/bridging-spec.txt	Mon Jan 28 16:55:42 2008 -0500
+++ b/bridging-spec.txt	Tue Jan 29 11:04:44 2008 -0500
@@ -235,12 +235,36 @@ 1.1 New dladm subcommands
 
       The options are the same as for the "create-bridge" subcommand.
 
-    dladm show-bridge [-p] [-s [-i <interval>]] [<bridge-name>]
+    dladm show-bridge [-p] [-o field,...] [-s [-i <interval>]]
+      [<bridge-name>]
 
       This subcommand shows the running status of bridges.  When given
       a bridge name, it shows the status of that one bridge.  If no
       bridge name is given, then it shows summary status of all
       bridges on the system.
+
+      The '-o' option allows the user to specify a comma-separated
+      case-insensitive list of fields to display.  The field name may
+      "all" to display all fields, or any combination of:
+
+	BRIDGE		Assigned name of the bridge (same as
+			<bridge-name>, if provided)
+	BRIDGEID	Bridge Identifier value (MAC + priority)
+	PRIORITY	Configured priority value (-p)
+	BMAXAGE		Configured bridge maximum age (-m)
+	BHELLOTIME	Configured bridge hello time (-h)
+	BFWDDELAY	Configured forwarding delay (-d)
+	FORCEPROTO	Configured forced maximum protocol (-f)
+	TCTIME		Time since last topology change in seconds
+	TCCOUNT		Count of the number of topology changes
+	TCHANGE		Topology change detected ("yes" or "no")
+	DESROOT		Bridge Identifier of the root node (MAC + priority)
+	ROOTCOST	Cost of the path to the root node
+	ROOTPORT	Port used to reach root node
+	MAXAGE		Maximum age value from root node
+	HELLOTIME	Hello time value from root node
+	FWDDELAY	Forward delay value from root node
+	HOLDTIME	Minimum BPDU interval
 
       Note the lack of a "-R" option here.  It is not possible to list
       bridge configuration information in an alternate root, in
@@ -251,7 +275,30 @@ 1.1 New dladm subcommands
       but "reading" is not feasible because the repository on the
       alternate root may be incompatible with the running system.
 
+    dladm show-bridge -P [-p] [-o field,...] [-s [-i <interval>]]
+      <bridge-name>
+
+      This variant of the show-bridge subcommand displays port-related
+      information for a single bridge instance.  Note that configured
+      parameters are shown through show-linkprop.  The relevant field
+      names for the "show-bridge -P" subcommand are:
+
+	PORT		Link name
+	STATE		"disabled", "listening", "learning",
+			"forwarding", or "blocking"
+	UPTIME		Number of seconds since last reset or initialize
+	DESROOT		Root Bridge Identifier (MAC + priority) seen
+			on this port
+	DESCOST		Path cost to root node through designated port
+	DESBRIDGE	Bridge Identifier (MAC + priority)
+	DESPORT		Port ID and priority of port used to transmit
+			configuration messages for this port
+	TCACK		Topology Change Acknowledge flag ("yes" or "no")
+
 1.2 New dladm Link Properties
+
+    These may be used with the existing dladm set-linkprop,
+    reset-linkprop, and show-linkprop subcommands.
 
     "stp"
 
@@ -590,6 +637,7 @@ 6.  Interface Summary
     Interface		Stability		Comments
     ---------		---------		--------
     dladm *-bridge	Committed
+    field names		Committed		dladm show-bridge -o
     link properties	Committed
     kstats		Volatile		Should be raised later
     /dev/bridge/	Committed		Observability node
@@ -599,3 +647,6 @@ 6.  Interface Summary
     config/*		Project Private		SMF properties
     bridge module	Project Private		Kernel bridging module
     DLPI_BRIDGE		Committed		dlpi_open(3DLPI)
+    /var/run/bridge_door/
+			Project Private		Doors interface to daemons
+    librstp.so.1	Project Private		RSTP implementation

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From erik.nordmark@sun.com Tue Jan 29 10:29:48 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0TITmge022603
	for <psarc-ext@sac.sfbay.sun.com>; Tue, 29 Jan 2008 10:29:48 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m0TITl5H042110;
	Tue, 29 Jan 2008 11:29:47 -0700 (MST)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVF00B094PLGA00@nwk-avmta-1.sfbay.Sun.COM>; Tue,
 29 Jan 2008 10:29:45 -0800 (PST)
Received: from jurassic.eng.sun.com ([129.146.228.31])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVF00AK64PJOK10@nwk-avmta-1.sfbay.Sun.COM>; Tue,
 29 Jan 2008 10:29:43 -0800 (PST)
Received: from [10.7.251.248] (punchin-nordmark.SFBay.Sun.COM [10.7.251.248])
	by jurassic.eng.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0TITgGW859825
	(version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-SHA bits=256 verify=NO); Tue,
 29 Jan 2008 10:29:43 -0800 (PST)
Date: Tue, 29 Jan 2008 10:29:42 -0800
From: Erik Nordmark <erik.nordmark@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <18334.20484.557823.2174@gargle.gargle.HOWL>
To: James Carlson <james.d.carlson@sun.com>
Cc: psarc-ext@sun.com
Message-id: <479F7096.9010405@sun.com>
MIME-version: 1.0
Content-type: text/plain; charset=ISO-8859-1; format=flowed
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.4 (X11/20070723)
Status: RO
Content-Length: 4025

James Carlson wrote:
> I'm sponsoring this fast-track request for myself.  The timer is set
> to 02/04/2008.  The release binding is "Minor" (because we depend on
> projects that have Minor binding and do not plan to work on a
> back-port) and the interface stability is listed at the end of the
> document.

Overall I'm very happy with this specification. I only have a few nits 
and questions below.

> The bridging protocol referred to in this document is the IEEE
> 802.1D-1998 "Spanning Tree Protocol," abbreviated in this document as
> "STP."  The newer and far more complex "Multiple Spanning Tree
> Protocol" (802.1Q-2005; MSTP) is intended to be backward compatible
> with STP, and is not part of this project, but may be the subject of a
> future project.

OK


>     dladm create-bridge [-t] [-R <root-dir>] [-p <priority>]
>       [-m <max-age>] [-h <hello-time>] [-d <forward-delay>]
>       [-f <force-protocol>] [-l <link>]... <bridge-name>

Given that there might be some future project that does RSTP or MSTP, 
wouldn't it make sense to have some notion of "bridge type" in the 
create-bridge syntax? That would seem to make it easier to introduce new 
types down the road. Then the initial type can be called "STP" and/or 
"802.1D-1998".

If the create-bridge is supposed to be limited to the backwards 
compatible set of IEEE 802.1D/802.1Q bridge protocols then you probably 
don't need a type since '-f' can be used to control those. Do you intend 
to limit it in that way?

>       This command creates a bridge instance and optionally assigns
>       network links to the new bridge.  By default, no bridge
>       instances are present, and OpenSolaris will not bridge between
>       network links.  See the "add-bridge" subcommand for details on
>       link assignment.

You say "assigns links". Does that mean that there can be multiple -l 
options to specify multiple links? Or is this subcommand limited to at 
most one link?


>       In this initial version, the links must also be Ethernet type.

Do we have a well-define notion of "Ethernet" for dladm?
Typing various dladm show-* on my machine just seems to tell me that my 
links are "non-vlan", which isn't sufficient to tell e.g., Ethernet and 
802.11 apart.


> 1.2 New dladm Link Properties
> 
>     "stp"
> 
> 	This is a boolean property.  It defaults to "true."  When set
> 	to "false," the link will not use Spanning Tree, and will be
> 	placed into forwarding mode at all times.  The "false" setting
> 	is appropriate for point-to-point links connected to end
> 	nodes.  Only non-VLAN type links have this property.

I wonder if it would be better, since we might have rstp or mstp in the 
future, to use Cisco type terminology for this (which I think is 
"portfast" or something like that.)

>     "stp-priority"
>     "stp-cost"

Could we/should we avoid "stp" in those names?

>     By default, all ports run standard STP.  This is done for safety
>     reasons: a bridge that does not run some form of bridging protocol
>     (such as STP) can form long-lasting forwarding loops in the
>     network.  Because Ethernet has no hop-count or TTL on packets, any
>     such loops are fatal to the network.
> 
>     When the adminstrator knows that a particular port is not
>     connected to another bridge (for example, a direct point-to-point
>     connection to a host system), STP can be disabled administratively
>     for that port.  Even if all ports on a bridge have STP disabled,
>     the STP daemon still runs; this is in case new ports are added,
>     and because it is responsible for enabling and disabling
>     forwarding on the ports.

This is a design issue and not an architectural issue. (Thus I should 
really have sent it as a separate email to the project team.):
For safety reasons, does the daemon watch for Bridge PDUs on the ports 
that have STP disabled? (There shouldn't be any, thus there existence is 
a red flag that the link probably isn't a point-to-point connection to a 
host system.)

    Erik

From erik.nordmark@sun.com Tue Jan 29 10:36:42 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0TIagj7022778
	for <psarc-ext@sac.sfbay.sun.com>; Tue, 29 Jan 2008 10:36:42 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m0TIad4w044447;
	Tue, 29 Jan 2008 11:36:40 -0700 (MST)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVF0040H513JX00@nwk-avmta-2.sfbay.sun.com>; Tue,
 29 Jan 2008 10:36:39 -0800 (PST)
Received: from jurassic.eng.sun.com ([129.146.226.31])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVF00M7Q511L630@nwk-avmta-2.sfbay.sun.com>; Tue,
 29 Jan 2008 10:36:37 -0800 (PST)
Received: from [10.7.251.248] (punchin-nordmark.SFBay.Sun.COM [10.7.251.248])
	by jurassic.eng.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0TIaGDM859929
	(version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-SHA bits=256 verify=NO); Tue,
 29 Jan 2008 10:36:26 -0800 (PST)
Date: Tue, 29 Jan 2008 10:36:16 -0800
From: Erik Nordmark <erik.nordmark@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <401216.62773.qm@web30801.mail.mud.yahoo.com>
To: Octave Orgeron <unixconsole@yahoo.com>
Cc: James Carlson <james.d.carlson@sun.com>, psarc-ext@sun.com
Message-id: <479F7220.6000001@sun.com>
MIME-version: 1.0
Content-type: text/plain; charset=ISO-8859-1; format=flowed
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <401216.62773.qm@web30801.mail.mud.yahoo.com>
User-Agent: Thunderbird 2.0.0.4 (X11/20070723)
Status: RO
Content-Length: 847

Octave Orgeron wrote:
> Hi James,
> 
> It's good to see this come up for integration. I was wondering how it'll impact virtualization (containers, xen, and ldoms)? Will further work be required or will it work transparently with things like VSW's in LDoms?

In addition to Jim's comments it makes sense to add that Xen and LDOMs 
don't require bridging support in Solaris.

In the control domain (Dom0 if you'd like) we use VNICs instead - 
basically handing out a fraction of a NIC to a domU (where the fraction 
is a MAC address today, but will be augmented with bandwidth limits by 
project Crossbow.

The VNIC abstraction makes it a lot more natural to expose the NICs 
characteristics and capabilities (hardware checksum, LSO, etc) to the 
domUs, than the current Linux approach of using a bridge to connect the 
domUs to the NICs.

    Erik

From carlsonj@phorcys.east.sun.com Tue Jan 29 10:58:47 2008
Received: from sunmail2sca.sfbay.sun.com (sunmail2sca [129.145.155.234])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0TIwlpI023800
	for <psarc-ext@sac.sfbay.sun.com>; Tue, 29 Jan 2008 10:58:47 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail2sca.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m0TIwkJD023767;
	Tue, 29 Jan 2008 10:58:46 -0800 (PST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVF0010B61Y8K00@brm-avmta-1.central.sun.com>; Tue,
 29 Jan 2008 11:58:46 -0700 (MST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVF007CQ61XBIF0@brm-avmta-1.central.sun.com>; Tue,
 29 Jan 2008 11:58:45 -0700 (MST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m0TIwjkW009906; Tue,
 29 Jan 2008 13:58:45 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m0TIwj0D009903; Tue,
 29 Jan 2008 13:58:45 -0500 (EST)
Date: Tue, 29 Jan 2008 13:58:45 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <479F7096.9010405@sun.com>
To: Erik Nordmark <Erik.Nordmark@sun.com>
Cc: psarc-ext@sun.com
Message-id: <18335.30565.241769.486286@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <479F7096.9010405@sun.com>
Status: RO
Content-Length: 5427

Erik Nordmark writes:
> James Carlson wrote:
> >     dladm create-bridge [-t] [-R <root-dir>] [-p <priority>]
> >       [-m <max-age>] [-h <hello-time>] [-d <forward-delay>]
> >       [-f <force-protocol>] [-l <link>]... <bridge-name>
> 
> Given that there might be some future project that does RSTP or MSTP, 

The plan is to support RSTP in this project because that's what the
code base I've got supports.  ;-}  Some future project may support
MSTP.

> wouldn't it make sense to have some notion of "bridge type" in the 
> create-bridge syntax? That would seem to make it easier to introduce new 
> types down the road. Then the initial type can be called "STP" and/or 
> "802.1D-1998".

The intent here is to allow for upgrade via a "natural" path.

In other words, when (and if) we support MSTP, an existing system that
has bridges configured and is upgraded to that new software will
automatically start running MSTP -- unless the user explicitly
disables such an upgrade by way of the IEEE-specified "force protocol"
feature.

> If the create-bridge is supposed to be limited to the backwards 
> compatible set of IEEE 802.1D/802.1Q bridge protocols then you probably 
> don't need a type since '-f' can be used to control those. Do you intend 
> to limit it in that way?

Create-bridge is supposed to be limited to a compatible set of
bridging protocols, so a "type" shouldn't be needed.

My current plan for the future RBridges/TRILL support (if that's what
you're actually asking about here) is to add parameters that allow
per-bridge disabling of RBridge behavior along with explicit
administrative action (via SMF) to enable a TRILL instance of isisd.
I could also be argued into a "bridge-created == isisd-enabled"
position for ease-of-use, but that's not this case.

The two protocols need to run side-by-side, at least by default.  (If
that's the other question here.  ;-})

> >       This command creates a bridge instance and optionally assigns
> >       network links to the new bridge.  By default, no bridge
> >       instances are present, and OpenSolaris will not bridge between
> >       network links.  See the "add-bridge" subcommand for details on
> >       link assignment.
> 
> You say "assigns links". Does that mean that there can be multiple -l 
> options to specify multiple links? Or is this subcommand limited to at 
> most one link?

The summary includes "[-l <link>]...", which means that (like the
create-aggr command on which it's modeled) you can add multiple links
at the time the bridge is created.

You can also add and remove one or many links after the bridge has
been created, using the add-bridge and remove-bridge subcommands,
again just like the corresponding *-aggr subcommands.

> >       In this initial version, the links must also be Ethernet type.
> 
> Do we have a well-define notion of "Ethernet" for dladm?
> Typing various dladm show-* on my machine just seems to tell me that my 
> links are "non-vlan", which isn't sufficient to tell e.g., Ethernet and 
> 802.11 apart.

The set of acceptable links is the same set that the *-aggr feature
supports, plus aggregations themselves.

Maybe I'm missing something important here, but I don't see a
particular reason to distinguish between 802.11 and other kinds of
interfaces.

> > 1.2 New dladm Link Properties
> > 
> >     "stp"
> > 
> > 	This is a boolean property.  It defaults to "true."  When set
> > 	to "false," the link will not use Spanning Tree, and will be
> > 	placed into forwarding mode at all times.  The "false" setting
> > 	is appropriate for point-to-point links connected to end
> > 	nodes.  Only non-VLAN type links have this property.
> 
> I wonder if it would be better, since we might have rstp or mstp in the 
> future, to use Cisco type terminology for this (which I think is 
> "portfast" or something like that.)

I'm very wary of and I don't want to step in any of their IPR.  :-/

> >     "stp-priority"
> >     "stp-cost"
> 
> Could we/should we avoid "stp" in those names?

The values are defined in terms that STP itself uses, and the IS-IS
notion of cost and priority are rather different, so having some sort
of reference to the domain in which the values are used makes sense to
me.

If we removed that term from those names, what would the units be?

> >     When the adminstrator knows that a particular port is not
> >     connected to another bridge (for example, a direct point-to-point
> >     connection to a host system), STP can be disabled administratively
> >     for that port.  Even if all ports on a bridge have STP disabled,
> >     the STP daemon still runs; this is in case new ports are added,
> >     and because it is responsible for enabling and disabling
> >     forwarding on the ports.
> 
> This is a design issue and not an architectural issue. (Thus I should 
> really have sent it as a separate email to the project team.):
> For safety reasons, does the daemon watch for Bridge PDUs on the ports 
> that have STP disabled? (There shouldn't be any, thus there existence is 
> a red flag that the link probably isn't a point-to-point connection to a 
> host system.)

That sounds like a reasonable thing to do.  I'll add it to the list.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From carlsonj@phorcys.east.sun.com Tue Jan 29 12:21:46 2008
Received: from sunmail3mpk.sfbay.sun.com (sunmail3mpk [129.146.11.52])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0TKLkl9026282
	for <psarc-ext@sac.sfbay.sun.com>; Tue, 29 Jan 2008 12:21:46 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail3mpk.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m0TKLdvK025901;
	Tue, 29 Jan 2008 12:21:46 -0800 (PST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVF0070H9W9ZF00@brm-avmta-1.central.sun.com>; Tue,
 29 Jan 2008 13:21:45 -0700 (MST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVF001NO9UOCB50@brm-avmta-1.central.sun.com>; Tue,
 29 Jan 2008 13:20:48 -0700 (MST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m0TKKmuD010360; Tue,
 29 Jan 2008 15:20:48 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m0TKKm2F010357; Tue,
 29 Jan 2008 15:20:48 -0500 (EST)
Date: Tue, 29 Jan 2008 15:20:48 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <479F7220.6000001@sun.com>
To: Erik Nordmark <Erik.Nordmark@sun.com>
Cc: Octave Orgeron <unixconsole@yahoo.com>, psarc-ext@sun.com
Message-id: <18335.35488.178884.844226@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <401216.62773.qm@web30801.mail.mud.yahoo.com>
 <479F7220.6000001@sun.com>
Status: RO
Content-Length: 1368

Erik Nordmark writes:
> The VNIC abstraction makes it a lot more natural to expose the NICs 
> characteristics and capabilities (hardware checksum, LSO, etc) to the 
> domUs, than the current Linux approach of using a bridge to connect the 
> domUs to the NICs.

Right; I agree with that.  If that's the purpose of using bridges in
Xen/LDOMs, then real 802-type bridges aren't what you want, and the
VNIC abstraction is what you need.

If the purpose is just to create a separate OS instance to run the
bridging software (because you don't trust the daemons, perhaps), then
running this new feature in Xen or an LDOM makes more sense.

For what it's worth, this bridging project is about constructing
802-type bridges with Solaris, which means taking packets in one
physical interface and forwarding them out another.  It faces
"downward" towards the interfaces.

Other quasi-bridge-like things are out of scope, and many of them
(such as the cases you're citing where packets are "bridged" between
virtual nodes) are better handled by VNICs.  The sort of learning and
loop prevention mechanisms required for regular bridges don't apply
there.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Darren.Reed@sun.com Tue Jan 29 23:17:35 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0U7HYd8011558
	for <psarc-ext@sac.sfbay.sun.com>; Tue, 29 Jan 2008 23:17:34 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id m0U7HUgh010766
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Wed, 30 Jan 2008 07:17:33 GMT
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVG00L09497DU00@nwk-avmta-1.sfbay.Sun.COM> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Tue, 29 Jan 2008 23:17:31 -0800 (PST)
Received: from sineb-mail-2.sun.com ([192.18.19.7])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVG007SN4965V80@nwk-avmta-1.sfbay.Sun.COM> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Tue,
 29 Jan 2008 23:17:31 -0800 (PST)
Received: from fe-apac-03.sun.com
 (fe-apac-03.sun.com [192.18.19.174] (may be forged))
	by sineb-mail-2.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m0U7HVNA023763	for
 <psarc-ext@sun.com>; Wed, 30 Jan 2008 07:17:31 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JVG00G013CQ5J00@mail-apac.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Wed,
 30 Jan 2008 15:17:29 +0800 (SGT)
Received: from [129.158.87.228] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JVG00AZO494TQH2@mail-apac.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Wed, 30 Jan 2008 15:17:29 +0800 (SGT)
Date: Wed, 30 Jan 2008 18:16:19 +1100
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <18334.20484.557823.2174@gargle.gargle.HOWL>
Sender: Darren.Reed@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: psarc-ext@sun.com
Message-id: <47A02443.5020207@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
User-Agent: Thunderbird 1.5.0.13 (Windows/20070809)
Status: RO
Content-Length: 1896

James Carlson wrote:
> ...
>     dladm show-bridge [-p] [-s [-i <interval>]] [<bridge-name>]
>
>       This subcommand shows the running status of bridges.  When given
>       a bridge name, it shows the status of that one bridge.  If no
>       bridge name is given, then it shows summary status of all
>       bridges on the system.
>
>       Note the lack of a "-R" option here.  It is not possible to list
>       bridge configuration information in an alternate root, in
>       keeping with the rest of the dladm user interface.  The reason
>       for this restriction is to allow the data to be represented in
>       SMF, where "writing" to an alternate root is supported by way of
>       copying appropriate commands to $ROOT/var/svc/profile/upgrade,
>       but "reading" is not feasible because the repository on the
>       alternate root may be incompatible with the running system.

If there is no separate command, how do we observe
the state of the bridging tables in use internally,
as we might routing information with "netstat -r"?

> ...
> 2.  Packet Observability
> ...
>     The observability node is intended for use with snoop and
>     wireshark.  It behaves as a standard Ethernet interface, but does
>     not permit the transmission of packets.  All transmitted packets
>     are silently dropped.

Was there any consideration given to there being
a kstat counter to record this drop events?

Some other questions....
What is the rights profile for bridging?

Can I bridge interfaces with different MTUs?

If so, what happens to the larger packets?

Is the assignment of a NIC to a bridge mutually
exclusive with other use, say with plumbing an
IP interface on it?

Are there any interaction issues with bridging
and the default setting of the eeprom setting
"local-mac-address?"? (OOB setting is "false",
potentially giving all NICs the same MAC address.)

Darren


From carlsonj@phorcys.east.sun.com Wed Jan 30 06:02:40 2008
Received: from sunmail2sca.sfbay.sun.com (sunmail2sca [129.145.155.234])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0UE2et3019552
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 30 Jan 2008 06:02:40 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail2sca.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m0UE2bOd024125;
	Wed, 30 Jan 2008 06:02:38 -0800 (PST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVG0020BN0BEB00@brm-avmta-1.central.sun.com>; Wed,
 30 Jan 2008 07:02:35 -0700 (MST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVG004P7N0AOHC0@brm-avmta-1.central.sun.com>; Wed,
 30 Jan 2008 07:02:35 -0700 (MST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m0UE2JeW013203; Wed,
 30 Jan 2008 09:02:25 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m0UE2DqQ013200; Wed,
 30 Jan 2008 09:02:13 -0500 (EST)
Date: Wed, 30 Jan 2008 09:02:13 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <47A02443.5020207@Sun.COM>
To: Darren Reed <Darren.Reed@sun.com>
Cc: psarc-ext@sun.com
Message-id: <18336.33637.339541.527190@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <47A02443.5020207@Sun.COM>
Status: RO
Content-Length: 3785

Darren Reed writes:
> James Carlson wrote:
> >       Note the lack of a "-R" option here.  It is not possible to list
> >       bridge configuration information in an alternate root, in
> >       keeping with the rest of the dladm user interface.  The reason
> >       for this restriction is to allow the data to be represented in
> >       SMF, where "writing" to an alternate root is supported by way of
> >       copying appropriate commands to $ROOT/var/svc/profile/upgrade,
> >       but "reading" is not feasible because the repository on the
> >       alternate root may be incompatible with the running system.
> 
> If there is no separate command, how do we observe
> the state of the bridging tables in use internally,
> as we might routing information with "netstat -r"?

Good question.  I think enhancing netstat is probably the best way to
do this.  BSD uses protocol 'bdg' to display bridging statistics, and
given the prior art, we should do likewise.

> >     The observability node is intended for use with snoop and
> >     wireshark.  It behaves as a standard Ethernet interface, but does
> >     not permit the transmission of packets.  All transmitted packets
> >     are silently dropped.
> 
> Was there any consideration given to there being
> a kstat counter to record this drop events?

No.  Trying to transmit on one of the /dev/bridge/* nodes would be a
fairly difficult blunder to wander into, so I don't see much point in
doing that.

It's certainly something that could be added later, though, if someone
felt it was necessary.  The kstat layout in section 1.3 would
accomodate it.

> Some other questions....
> What is the rights profile for bridging?

No new rights profile or change to existing profiles is needed.  The
existing "Network Link Security" and "Network Management" rights
profiles include dladm with sufficient privilege (as documented in
this project) to allow administration of bridges.

> Can I bridge interfaces with different MTUs?

It's permitted per the standards, but a foolish thing to do.  We
should emit a warning in this case.

A future project may implement LLDP to help prevent toxic mixed MTUs
on a LAN.

> If so, what happens to the larger packets?

They're silently dropped, per 802.1Q-2005 section 6.3.8.  (Yes, we can
have a kstat for this.)

> Is the assignment of a NIC to a bridge mutually
> exclusive with other use, say with plumbing an
> IP interface on it?

No.

To get more detail on that answer, we get into detailed design issues.
For instance, on transmit of an IP packet through an interface
attached to a bridge, we must do a bridge forwarding table look-up on
the MAC destination address, and forward it as though it were received
on that interface, except that if the forwarding entry points back to
the interface, we transmit rather than drop.

> Are there any interaction issues with bridging
> and the default setting of the eeprom setting
> "local-mac-address?"? (OOB setting is "false",
> potentially giving all NICs the same MAC address.)

That's not quite true.  The out-of-the-box setting for
"local-mac-address?" has been "true" for years -- see CRs 4473325,
4798630, and 6504375, among others.

There are of course some hold-outs, as in CR 4868371.  

You're right that the "local-mac-address?=false" misfeature is a
tremendous call generator, as it breaks at least IPMP, DHCP, and RARP
and often breaks other things as well.  Given how well-known the
problem is, and the trend towards fixing this, I'm not positive that
documentation is needed, but I'll add it anyway.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From erik.nordmark@sun.com Wed Jan 30 13:32:34 2008
Received: from sunmail2sca.sfbay.sun.com (sunmail2sca [129.145.155.234])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0ULWYH6019586
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 30 Jan 2008 13:32:34 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail2sca.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m0ULWViY005430;
	Wed, 30 Jan 2008 13:32:33 -0800 (PST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVH00D057U88S00@brm-avmta-1.central.sun.com>; Wed,
 30 Jan 2008 14:32:32 -0700 (MST)
Received: from jurassic.eng.sun.com ([129.146.106.105])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVH007BA7U7OC80@brm-avmta-1.central.sun.com>; Wed,
 30 Jan 2008 14:32:31 -0700 (MST)
Received: from [10.7.251.248] (punchin-nordmark.SFBay.Sun.COM [10.7.251.248])
	by jurassic.eng.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0ULWUT3928484
	(version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-SHA bits=256 verify=NO); Wed,
 30 Jan 2008 13:32:30 -0800 (PST)
Date: Wed, 30 Jan 2008 13:32:30 -0800
From: Erik Nordmark <erik.nordmark@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <18335.30565.241769.486286@gargle.gargle.HOWL>
To: James Carlson <james.d.carlson@sun.com>
Cc: psarc-ext@sun.com
Message-id: <47A0ECEE.7020808@sun.com>
MIME-version: 1.0
Content-type: text/plain; charset=ISO-8859-1; format=flowed
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <479F7096.9010405@sun.com> <18335.30565.241769.486286@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.4 (X11/20070723)
Status: RO
Content-Length: 4059

James Carlson wrote:

> Create-bridge is supposed to be limited to a compatible set of
> bridging protocols, so a "type" shouldn't be needed.
> 
> My current plan for the future RBridges/TRILL support (if that's what
> you're actually asking about here) is to add parameters that allow
> per-bridge disabling of RBridge behavior along with explicit
> administrative action (via SMF) to enable a TRILL instance of isisd.
> I could also be argued into a "bridge-created == isisd-enabled"
> position for ease-of-use, but that's not this case.

Well, rbridges we certainly on my mind.
And the internet-draft doesn't require an rbridge to embed full 802.1D/Q 
functionality - even though it probably makes sense for many 
implementations to do that.

Anyhow, you've convinced me we don't need a type.

> The summary includes "[-l <link>]...", which means that (like the
> create-aggr command on which it's modeled) you can add multiple links
> at the time the bridge is created.

Ah - I didn't understand that detail of the syntax.

> You can also add and remove one or many links after the bridge has
> been created, using the add-bridge and remove-bridge subcommands,
> again just like the corresponding *-aggr subcommands.

Yes, I saw that.

> The set of acceptable links is the same set that the *-aggr feature
> supports, plus aggregations themselves.
> 
> Maybe I'm missing something important here, but I don't see a
> particular reason to distinguish between 802.11 and other kinds of
> interfaces.

Bridges require two things: being able to promisciously receive frames, 
and being able to send frames with any source MAC address.
While 802.11 standards might not prevent those, it isn't uncommon for 
802.11 NICs (or drivers) to prevent using anything bit the assigned MAC 
address.

Hmm - perhaps there should be link property (send-why-any-source?) that 
can be read by the bridge code.

>> I wonder if it would be better, since we might have rstp or mstp in the 
>> future, to use Cisco type terminology for this (which I think is 
>> "portfast" or something like that.)
> 
> I'm very wary of and I don't want to step in any of their IPR.  :-/

OK, but we should at least make the description more clear to people 
that know and understand portfast.

But I realize they are not the same. portfast is described as
By enabling portfast you are forcing the switchport into forwarding mode
immediately.  The port still participates in STP in the event that if
the port is to be a part of the loop, it will eventually transition into
STP blocking mode.

Thus disabling 802.1D completely things are a lot less safe than just 
avoiding the 802.1D delay before being transitioning to forwarding.
See below for a name suggestion.

>>>     "stp-priority"
>>>     "stp-cost"
>> Could we/should we avoid "stp" in those names?
> 
> The values are defined in terms that STP itself uses, and the IS-IS
> notion of cost and priority are rather different, so having some sort
> of reference to the domain in which the values are used makes sense to
> me.
> 
> If we removed that term from those names, what would the units be?

My concern is that "stp" seems to refer to a very old version of a 
protocol, that has since been replaced by rstp and mstp.
Given that we have rstp, we don't want to look as if we have old stp.

FWIW "bridge-priority" and "bridge-cost" works for me.

>> This is a design issue and not an architectural issue. (Thus I should 
>> really have sent it as a separate email to the project team.):
>> For safety reasons, does the daemon watch for Bridge PDUs on the ports 
>> that have STP disabled? (There shouldn't be any, thus there existence is 
>> a red flag that the link probably isn't a point-to-point connection to a 
>> host system.)
> 
> That sounds like a reasonable thing to do.  I'll add it to the list.

I guess the interface question is whether we call it "disable stp" even 
though 802.1D is still active on the link, or whether we instead 
introduce a name like "bridge-immediate-forwarding" for this configuration.

    Erik



From carlsonj@phorcys.east.sun.com Wed Jan 30 14:11:07 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0UMB6RA023796
	for <psarc-ext@sac.sfbay.Sun.COM>; Wed, 30 Jan 2008 14:11:07 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id m0UMB3LY029335;
	Thu, 31 Jan 2008 06:11:03 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVH00M059MD2D00@nwk-avmta-2.sfbay.sun.com>; Wed,
 30 Jan 2008 14:11:01 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVH00G579MCHAA0@nwk-avmta-2.sfbay.sun.com>; Wed,
 30 Jan 2008 14:11:00 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m0UMAu46016270; Wed,
 30 Jan 2008 17:10:56 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m0UMAuIR016267; Wed,
 30 Jan 2008 17:10:56 -0500 (EST)
Date: Wed, 30 Jan 2008 17:10:56 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <47A0ECEE.7020808@sun.com>
To: Erik Nordmark <Erik.Nordmark@sun.com>
Cc: psarc-ext@sun.com
Message-id: <18336.62960.455810.182556@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <479F7096.9010405@sun.com> <18335.30565.241769.486286@gargle.gargle.HOWL>
 <47A0ECEE.7020808@sun.com>
Status: RO
Content-Length: 5884

Erik Nordmark writes:
> James Carlson wrote:
> > My current plan for the future RBridges/TRILL support (if that's what
> > you're actually asking about here) is to add parameters that allow
> > per-bridge disabling of RBridge behavior along with explicit
> > administrative action (via SMF) to enable a TRILL instance of isisd.
> > I could also be argued into a "bridge-created == isisd-enabled"
> > position for ease-of-use, but that's not this case.
> 
> Well, rbridges we certainly on my mind.
> And the internet-draft doesn't require an rbridge to embed full 802.1D/Q 
> functionality - even though it probably makes sense for many 
> implementations to do that.

I doubt implementations can really get away without it, but it does
make sense that the draft doesn't _require_ it.

> Anyhow, you've convinced me we don't need a type.

OK.

> > The set of acceptable links is the same set that the *-aggr feature
> > supports, plus aggregations themselves.
> > 
> > Maybe I'm missing something important here, but I don't see a
> > particular reason to distinguish between 802.11 and other kinds of
> > interfaces.
> 
> Bridges require two things: being able to promisciously receive frames, 
> and being able to send frames with any source MAC address.
> While 802.11 standards might not prevent those, it isn't uncommon for 
> 802.11 NICs (or drivers) to prevent using anything bit the assigned MAC 
> address.
> 
> Hmm - perhaps there should be link property (send-why-any-source?) that 
> can be read by the bridge code.

The very same issue occurs with aggregations, doesn't it?  You need to
be able to treat the aggregation as a single sender for the network
layer clients, which means sending from the "wrong" MAC address on at
least some of the links.

I don't think it's a new or bridging-specific issue.  If we have this
problem, then we've also got it with aggregations, which is part of
why I defined the usable links in those terms.  Our constraints on
link types are effectively the same as what aggregations can use, plus
aggregations themselves.

I think it gets outside the bounds of this case, but I agree that if
we have those sorts of drivers, we'll need the same logic to prevent
misconfiguration in both cases.  I'll follow up on the supported
drivers.

> >> I wonder if it would be better, since we might have rstp or mstp in the 
> >> future, to use Cisco type terminology for this (which I think is 
> >> "portfast" or something like that.)
> > 
> > I'm very wary of and I don't want to step in any of their IPR.  :-/
> 
> OK, but we should at least make the description more clear to people 
> that know and understand portfast.
> 
> But I realize they are not the same. portfast is described as
> By enabling portfast you are forcing the switchport into forwarding mode
> immediately.  The port still participates in STP in the event that if
> the port is to be a part of the loop, it will eventually transition into
> STP blocking mode.

It may ... after having blindly forwarded for long enough to melt the
network.

Or you can get into root-guard and BPDU-guard; more about that below.
The more I think about Cisco-like 'portfast', the less I'm sure we
want to support it in the first release.

> Thus disabling 802.1D completely things are a lot less safe than just 
> avoiding the 802.1D delay before being transitioning to forwarding.
> See below for a name suggestion.

OK.

> >>>     "stp-priority"
> >>>     "stp-cost"
> >> Could we/should we avoid "stp" in those names?
> > 
> > The values are defined in terms that STP itself uses, and the IS-IS
> > notion of cost and priority are rather different, so having some sort
> > of reference to the domain in which the values are used makes sense to
> > me.
> > 
> > If we removed that term from those names, what would the units be?
> 
> My concern is that "stp" seems to refer to a very old version of a 
> protocol, that has since been replaced by rstp and mstp.
> Given that we have rstp, we don't want to look as if we have old stp.

Ah, I see.  I was just using it as a generic term: RSTP and MSTP are
both instances of STP.  Just newer ones.

> FWIW "bridge-priority" and "bridge-cost" works for me.

OK; that seems fine.

> > That sounds like a reasonable thing to do.  I'll add it to the list.
> 
> I guess the interface question is whether we call it "disable stp" even 
> though 802.1D is still active on the link, or whether we instead 
> introduce a name like "bridge-immediate-forwarding" for this configuration.

I'm wary of this.  There's a bit more to it than just switching into
forwarding mode immediately, because that still allows a node to
attach via that port, send BPDUs with a lower (better) priority, and
become the root bridge in your network.  Given how some people are
forced to dance around the root bridge determination fire with severed
chicken heads, that's hazardous ... just in a different way.

If simply allowing an STP per-port disable seems too hazardous, then
I'd suggest we go with something like BPDU-guard.  We can still
disable STP, but if a BPDU is seen arriving on such a port, then we
disable the port as apparently broken (changing state to 'disabled')
and log a message.  The administrator must explicitly reenable the
link by removing it from the bridge and adding it back or otherwise
making the link signal go down and back up.  Seeing a BPDU on such a
port is good evidence that the topology is just plain wrong, and
saving the network at the expense of a port sounds like the right
result to me.

With, of course, warnings about the implications of doing this
included in the man pages and administrator's guide updates.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From erik.nordmark@sun.com Wed Jan 30 17:30:25 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0V1UOLo001110
	for <psarc-ext@sac.sfbay.Sun.COM>; Wed, 30 Jan 2008 17:30:24 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id m0V1UC7q012608;
	Thu, 31 Jan 2008 09:30:19 +0800 (SGT)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVH0090DIUI8C00@brm-avmta-1.central.sun.com>; Wed,
 30 Jan 2008 18:30:18 -0700 (MST)
Received: from jurassic.eng.sun.com ([129.146.104.31])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVH00I5MIUH8Z90@brm-avmta-1.central.sun.com>; Wed,
 30 Jan 2008 18:30:18 -0700 (MST)
Received: from [10.7.251.248] (punchin-nordmark.SFBay.Sun.COM [10.7.251.248])
	by jurassic.eng.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0V1UDQP933553
	(version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-SHA bits=256 verify=NO); Wed,
 30 Jan 2008 17:30:13 -0800 (PST)
Date: Wed, 30 Jan 2008 17:30:08 -0800
From: Erik Nordmark <erik.nordmark@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <18336.62960.455810.182556@gargle.gargle.HOWL>
To: James Carlson <james.d.carlson@sun.com>
Cc: psarc-ext@sun.com
Message-id: <47A124A0.9000808@sun.com>
MIME-version: 1.0
Content-type: text/plain; charset=ISO-8859-1; format=flowed
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <479F7096.9010405@sun.com> <18335.30565.241769.486286@gargle.gargle.HOWL>
 <47A0ECEE.7020808@sun.com> <18336.62960.455810.182556@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.4 (X11/20070723)
Status: RO
Content-Length: 1917

James Carlson wrote:
>
> The very same issue occurs with aggregations, doesn't it?  You need to
> be able to treat the aggregation as a single sender for the network
> layer clients, which means sending from the "wrong" MAC address on at
> least some of the links.

Link aggregation is IEEE 802.3AD i.e. defined only for IEEE 802.3.
I don't know how aggr currently checks that it is only talking to 802.3, 
but I'm assuming there is such a check.

Since bridging is an IEEE 802.1 standard in principle it applies to 
802.11 etc.

> I think it gets outside the bounds of this case, but I agree that if
> we have those sorts of drivers, we'll need the same logic to prevent
> misconfiguration in both cases.  I'll follow up on the supported
> drivers.

Good (but the logic check might be slightly different due to the above 
difference.)

> It may ... after having blindly forwarded for long enough to melt the
> network.
> 
> Or you can get into root-guard and BPDU-guard; more about that below.
> The more I think about Cisco-like 'portfast', the less I'm sure we
> want to support it in the first release.

But providing a way to disable stp completely is then even more 
dangerous thus I'd propose not to offer that in the first release either.

> If simply allowing an STP per-port disable seems too hazardous, then
> I'd suggest we go with something like BPDU-guard.  We can still
> disable STP, but if a BPDU is seen arriving on such a port, then we
> disable the port as apparently broken (changing state to 'disabled')
> and log a message.  The administrator must explicitly reenable the
> link by removing it from the bridge and adding it back or otherwise
> making the link signal go down and back up.  Seeing a BPDU on such a
> port is good evidence that the topology is just plain wrong, and
> saving the network at the expense of a port sounds like the right
> result to me.

That works for me.

    Erik

From Darren.Reed@sun.com Wed Jan 30 18:40:37 2008
Received: from sunmail3mpk.sfbay.sun.com (sunmail3mpk [129.146.11.52])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m0V2ebBg003385
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 30 Jan 2008 18:40:37 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail3mpk.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m0V2eYAN025205
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Wed, 30 Jan 2008 18:40:35 -0800 (PST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVH00F01M3NB900@brm-avmta-1.central.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@Sun.COM); Wed, 30 Jan 2008 19:40:35 -0700 (MST)
Received: from sineb-mail-1.sun.com ([192.18.19.6])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVH00IZBM3L98A0@brm-avmta-1.central.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@Sun.COM); Wed,
 30 Jan 2008 19:40:34 -0700 (MST)
Received: from fe-apac-01.sun.com
 (fe-apac-01.sun.com [192.18.19.172] (may be forged))
	by sineb-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m0V2eWCo018095	for
 <psarc-ext@Sun.COM>; Thu, 31 Jan 2008 02:40:32 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JVH00L01LHM6100@mail-apac.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@Sun.COM (ORCPT psarc-ext@Sun.COM); Thu,
 31 Jan 2008 10:40:32 +0800 (SGT)
Received: from [129.158.87.228] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JVH00GNVM3JVBF2@mail-apac.sun.com> for psarc-ext@Sun.COM
 (ORCPT psarc-ext@Sun.COM); Thu, 31 Jan 2008 10:40:32 +0800 (SGT)
Date: Thu, 31 Jan 2008 13:39:21 +1100
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <18336.33637.339541.527190@gargle.gargle.HOWL>
Sender: Darren.Reed@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: psarc-ext@sun.com
Message-id: <47A134D9.2080806@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <47A02443.5020207@Sun.COM> <18336.33637.339541.527190@gargle.gargle.HOWL>
User-Agent: Thunderbird 1.5.0.13 (Windows/20070809)
Status: RO
Content-Length: 524

James Carlson wrote:
> ...
>> Some other questions....
>> What is the rights profile for bridging?
>>     
>
> No new rights profile or change to existing profiles is needed.  The
> existing "Network Link Security" and "Network Management" rights
> profiles include dladm with sufficient privilege (as documented in
> this project) to allow administration of bridges


Will the daemon also be associated with one or both of these?

Is there a related authorisation profile for managing the bridging
SMF service(s)?

Darren


From Darren.Moffat@sun.com Mon Feb  4 07:20:26 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m14FKP4E002845
	for <psarc-ext@sac.sfbay.Sun.COM>; Mon, 4 Feb 2008 07:20:26 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id m14FK576027913
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Mon, 4 Feb 2008 23:20:24 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVP00I07ZXXCL00@nwk-avmta-2.sfbay.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Mon, 04 Feb 2008 07:20:22 -0800 (PST)
Received: from gmp-eb-mail-2.sun.com ([192.18.6.24])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVP00B78ZXQ7BB0@nwk-avmta-2.sfbay.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Mon,
 04 Feb 2008 07:20:15 -0800 (PST)
Received: from fe-emea-10.sun.com (gmp-eb-lb-2-fe3.eu.sun.com [192.18.6.12])
	by gmp-eb-mail-2.sun.com (8.13.7+Sun/8.12.9) with ESMTP id m14FKDnH022624	for
 <psarc-ext@sun.com>; Mon, 04 Feb 2008 15:20:13 +0000 (GMT)
Received: from conversion-daemon.fe-emea-10.sun.com by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0JVP00201YP4XU00@fe-emea-10.sun.com>
 (original mail from Darren.Moffat@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Mon,
 04 Feb 2008 15:20:13 +0000 (GMT)
Received: from [129.156.173.21] by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0JVP009TFZX15R10@fe-emea-10.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Mon, 04 Feb 2008 15:19:50 +0000 (GMT)
Date: Mon, 04 Feb 2008 15:19:49 +0000
From: Darren J Moffat <Darren.Moffat@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <47A134D9.2080806@Sun.COM>
Sender: Darren.Moffat@sun.com
To: Darren Reed <Darren.Reed@sun.com>
Cc: James Carlson <James.D.Carlson@sun.com>, psarc-ext@sun.com
Message-id: <47A72D15.7090708@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <47A02443.5020207@Sun.COM> <18336.33637.339541.527190@gargle.gargle.HOWL>
 <47A134D9.2080806@Sun.COM>
User-Agent: Thunderbird 2.0.0.6 (X11/20080102)
Status: RO
Content-Length: 1061

Darren Reed wrote:
> James Carlson wrote:
>> ...
>>> Some other questions....
>>> What is the rights profile for bridging?
>>>     
>>
>> No new rights profile or change to existing profiles is needed.  The
>> existing "Network Link Security" and "Network Management" rights
>> profiles include dladm with sufficient privilege (as documented in
>> this project) to allow administration of bridges
> 
> 
> Will the daemon also be associated with one or both of these?

Why should it be ?  The daemon should only be started by SMF.  While it 
is possible to write the SMF manifest such that it uses an exec_attr 
profile rather than explicit credential entries I don't think that is 
necessary.  In fact I'd say that unless the daemon is intended to also 
be started by a normal user (for something other than debug purposes) 
then using an RBAC profile in the SMF manifest just encourages users to 
think they can start the daemon manually (of course the daemon can be 
coded to check it is actually running under SMF and refuse to start!).

-- 
Darren J Moffat

From carlsonj@phorcys.east.sun.com Mon Feb  4 07:36:28 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m14FaRBq002943
	for <psarc-ext@sac.sfbay.sun.com>; Mon, 4 Feb 2008 07:36:28 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id m14FaA7U010709;
	Mon, 4 Feb 2008 15:36:23 GMT
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVQ0061J0OJTX00@brm-avmta-1.central.sun.com>; Mon,
 04 Feb 2008 08:36:19 -0700 (MST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVQ006080OETQ00@brm-avmta-1.central.sun.com>; Mon,
 04 Feb 2008 08:36:14 -0700 (MST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m14FaEIT008997; Mon,
 04 Feb 2008 10:36:14 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m14FaEZO008994; Mon,
 04 Feb 2008 10:36:14 -0500 (EST)
Date: Mon, 04 Feb 2008 10:36:14 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <47A72D15.7090708@Sun.COM>
To: Darren J Moffat <Darren.Moffat@sun.com>
Cc: Darren Reed <Darren.Reed@sun.com>, psarc-ext@sun.com
Message-id: <18343.12526.92706.814261@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <47A02443.5020207@Sun.COM> <18336.33637.339541.527190@gargle.gargle.HOWL>
 <47A134D9.2080806@Sun.COM> <47A72D15.7090708@Sun.COM>
Status: RO
Content-Length: 1642

Darren J Moffat writes:
> Darren Reed wrote:
> > James Carlson wrote:
> >> ...
> >>> Some other questions....
> >>> What is the rights profile for bridging?
> >>>     
> >>
> >> No new rights profile or change to existing profiles is needed.  The
> >> existing "Network Link Security" and "Network Management" rights
> >> profiles include dladm with sufficient privilege (as documented in
> >> this project) to allow administration of bridges
> > 
> > 
> > Will the daemon also be associated with one or both of these?
> 
> Why should it be ?  The daemon should only be started by SMF.  While it 
> is possible to write the SMF manifest such that it uses an exec_attr 
> profile rather than explicit credential entries I don't think that is 
> necessary.  In fact I'd say that unless the daemon is intended to also 
> be started by a normal user (for something other than debug purposes) 
> then using an RBAC profile in the SMF manifest just encourages users to 
> think they can start the daemon manually (of course the daemon can be 
> coded to check it is actually running under SMF and refuse to start!).

Exactly and, no, the user will not be expected to start the daemon
manually.  It requires SMF data to start correctly anyway.

(And, yes, the daemon will run with least privilege.)

I need to update the specification for this case, so I've placed it in
"waiting need spec" until I can draft a new document.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Darren.Reed@sun.com Tue Feb  5 21:23:32 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m165NVKd026657
	for <psarc-ext@sac.sfbay.Sun.COM>; Tue, 5 Feb 2008 21:23:31 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id m165N3nb017933
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Wed, 6 Feb 2008 13:23:30 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JVS0061PXN3KE00@nwk-avmta-1.sfbay.Sun.COM> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Tue, 05 Feb 2008 21:23:27 -0800 (PST)
Received: from sineb-mail-1.sun.com ([192.18.19.6])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JVS001R4XN1KF90@nwk-avmta-1.sfbay.Sun.COM> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Tue,
 05 Feb 2008 21:23:26 -0800 (PST)
Received: from fe-apac-05.sun.com
 (fe-apac-05.sun.com [192.18.19.176] (may be forged))
	by sineb-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m165NRKm004785	for
 <psarc-ext@sun.com>; Wed, 06 Feb 2008 05:23:27 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JVS00A01XI2GP00@mail-apac.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Wed,
 06 Feb 2008 13:23:25 +0800 (SGT)
Received: from [129.158.87.228] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JVS00353XN0MW2A@mail-apac.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Wed, 06 Feb 2008 13:23:25 +0800 (SGT)
Date: Wed, 06 Feb 2008 16:22:02 +1100
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: 2008/055 Solaris Bridging
In-reply-to: <18343.12526.92706.814261@gargle.gargle.HOWL>
Sender: Darren.Reed@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: Darren J Moffat <Darren.Moffat@sun.com>, psarc-ext@sun.com
Message-id: <47A943FA.1040002@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.2.0.264296
References: <18334.20484.557823.2174@gargle.gargle.HOWL>
 <47A02443.5020207@Sun.COM> <18336.33637.339541.527190@gargle.gargle.HOWL>
 <47A134D9.2080806@Sun.COM> <47A72D15.7090708@Sun.COM>
 <18343.12526.92706.814261@gargle.gargle.HOWL>
User-Agent: Thunderbird 1.5.0.13 (Windows/20070809)
Status: RO
Content-Length: 1909

James Carlson wrote:
> Darren J Moffat writes:
>   
>> Darren Reed wrote:
>>     
>>> James Carlson wrote:
>>>       
>>>> ...
>>>>         
>>>>> Some other questions....
>>>>> What is the rights profile for bridging?
>>>>>     
>>>>>           
>>>> No new rights profile or change to existing profiles is needed.  The
>>>> existing "Network Link Security" and "Network Management" rights
>>>> profiles include dladm with sufficient privilege (as documented in
>>>> this project) to allow administration of bridges
>>>>         
>>> Will the daemon also be associated with one or both of these?
>>>       
>> Why should it be ?  The daemon should only be started by SMF.  While it 
>> is possible to write the SMF manifest such that it uses an exec_attr 
>> profile rather than explicit credential entries I don't think that is 
>> necessary.  In fact I'd say that unless the daemon is intended to also 
>> be started by a normal user (for something other than debug purposes) 
>> then using an RBAC profile in the SMF manifest just encourages users to 
>> think they can start the daemon manually (of course the daemon can be 
>> coded to check it is actually running under SMF and refuse to start!).
>>     
>
> Exactly and, no, the user will not be expected to start the daemon
> manually.  It requires SMF data to start correctly anyway.
>
> (And, yes, the daemon will run with least privilege.)
>
> I need to update the specification for this case, so I've placed it in
> "waiting need spec" until I can draft a new document.
>   

I was actually going to let both of the questions in that email
of mine slide...I went back and did some more reading,
noticed that the daemon itself was a private interface (and
thus it isn't expected to be run manually by a user.)

The other question I asked was answered by Jim, I just
didn't plug everything together w.r.t profile names and
authorisations.

Darren


From sacadmin Tue Dec  9 08:00:22 2008
Received: from sunmail2sca.sfbay.sun.com (sunmail2sca.SFBay.Sun.COM [129.145.155.234])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mB9G0Mo0017803
	for <psarc@sac.eng.sun.com>; Tue, 9 Dec 2008 08:00:22 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail2sca.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mB9G0Ian026373
	for <@sunmail2sca.sfbay.sun.com:psarc@sun.com>; Tue, 9 Dec 2008 08:00:22 -0800 (PST)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBM00M0P9SL0800@nwk-avmta-2.sfbay.sun.com> for psarc@sun.com
 (ORCPT psarc@Sun.COM); Tue, 09 Dec 2008 08:00:21 -0800 (PST)
Received: from brmea-mail-1.sun.com ([192.18.98.31])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBM00HFG9SE4V80@nwk-avmta-2.sfbay.sun.com> for psarc@sun.com
 (ORCPT psarc@Sun.COM); Tue, 09 Dec 2008 08:00:14 -0800 (PST)
Received: from fe-amer-09.sun.com ([192.18.109.79])
	by brmea-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id mB9G0CNo010368	for
 <psarc@Sun.COM>; Tue, 09 Dec 2008 16:00:14 +0000 (GMT)
Received: from conversion-daemon.mail-amer.sun.com by mail-amer.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KBM004018KZ4900@mail-amer.sun.com> (original mail from Aarti.Pai@Sun.COM)
 for psarc@Sun.COM (ORCPT psarc@Sun.COM); Tue, 09 Dec 2008 09:00:13 -0700 (MST)
Received: from [129.145.154.95] by mail-amer.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KBM0070M9S3Y060@mail-amer.sun.com> for psarc@Sun.COM
 (ORCPT psarc@Sun.COM); Tue, 09 Dec 2008 09:00:04 -0700 (MST)
Date: Tue, 09 Dec 2008 08:00:03 -0800
From: Aarti Pai <Aarti.Pai@sun.com>
Subject: Re: 2008/055 materials
In-reply-to: <18750.31654.801877.584707@gargle.gargle.HOWL>
Sender: Aarti.Pai@sun.com
To: psarc@sun.com
Cc: James Carlson <James.D.Carlson@sun.com>
Reply-to: Aarti.Pai@sun.com
Message-id: <493E9603.6010803@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <18750.31654.801877.584707@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.16 (X11/20080807)
Status: RO
Content-Length: 916

sac% pwd
/shared/sac/Archives/CaseLog/arc/PSARC/2008/055/inception.materials
sac% ls -l
total 930
-rw-r--r--   1 carlsonj sac         8744 Dec  9 05:56 bridging-20q-new.txt
-rw-r--r--   1 carlsonj sac        21169 Dec  9 05:56 bridging-20q.txt
-rw-r--r--   1 carlsonj sac       136324 Dec  9 05:56 bridging-design.pdf
-rw-r--r--   1 carlsonj sac         8564 Dec  9 05:56 bridging-security.txt
-rw-r--r--   1 carlsonj sac        49747 Dec  9 05:57 bridging-spec.txt
-rw-r--r--   1 carlsonj sac       114292 May 21  2008 
rbridges-net-summit.odp
-rw-r--r--   1 carlsonj sac         1219 Dec  9 06:07 README.txt





On 12/ 9/08 06:07 AM, James Carlson wrote:
> I've deposited the inception review materials for next week into the
> case directory.
>
> I've included both new and old style 20q documents, as I wrote the 20q
> before the new document came out.  There's a README.txt that ties it
> all together.
>
>   

From Nicolas.Droux@sun.com Wed Dec 17 10:04:10 2008
Received: from sunmail3mpk.sfbay.sun.com (sunmail3mpk.SFBay.Sun.COM [129.146.11.52])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBHI4A5B007872
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 17 Dec 2008 10:04:10 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail3mpk.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mBHI46s2022360
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Wed, 17 Dec 2008 10:04:10 -0800 (PST)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KC100A0F8UXOG00@nwk-avmta-1.sfbay.Sun.COM> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Wed, 17 Dec 2008 10:04:09 -0800 (PST)
Received: from brmea-mail-4.sun.com ([192.18.98.36])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KC100H2W8UVKQB0@nwk-avmta-1.sfbay.Sun.COM> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Wed,
 17 Dec 2008 10:04:08 -0800 (PST)
Received: from fe-amer-10.sun.com ([192.18.109.80])
	by brmea-mail-4.sun.com (8.13.6+Sun/8.12.9) with ESMTP id mBHI47g2009100	for
 <psarc-ext@sun.com>; Wed, 17 Dec 2008 18:04:07 +0000 (GMT)
Received: from conversion-daemon.mail-amer.sun.com by mail-amer.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KC1000015T82E00@mail-amer.sun.com>
 (original mail from Nicolas.Droux@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Wed,
 17 Dec 2008 11:04:07 -0700 (MST)
Received: from [10.0.0.2] ([129.150.18.107])
 by mail-amer.sun.com (Sun Java System Messaging Server 6.2-8.04 (built Feb 28
 2007)) with ESMTPSA id <0KC1002JS8UNIB40@mail-amer.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Wed,
 17 Dec 2008 11:04:01 -0700 (MST)
Date: Wed, 17 Dec 2008 11:03:58 -0700
From: Nicolas Droux <Nicolas.Droux@sun.com>
Subject: Issues for 2008/055
Sender: Nicolas.Droux@sun.com
To: PSARC-ext@sun.com
Message-id: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
MIME-version: 1.0
X-Mailer: Apple Mail (2.930.3)
Content-type: text/plain; delsp=yes; format=flowed; charset=US-ASCII
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
Status: RO
Content-Length: 1974

Here are my issues for Solaris Bridging (PSARC 2008/055).  
Unfortunately I have a conflict and won't be able to attend the  
inception review, but I'll be happy to follow-up by email.

ngd-01 bridging-spec.txt states "The links assigned to a bridge must  
not themselves be VLANs, VNICs, or tunnels. Only links that would be  
acceptable as part of an aggregation or links that are aggregations  
themselves may be assigned to a bridge." It should be also possible to  
bridge etherstubs (introduced by Crossbow [PSARC 2006/357]), since  
they can be used to create virtual switches.

ngd-02 in bridging-spec.txt, 2.2 a), the proposed link/up behavior in  
the presence of bridges needs to be refined. With Crossbow VNICs, the  
link status advertised to MAC clients depends also on the presence of  
other MAC clients on top of the underlying data-link, in order to  
maintain connectivity between these MAC clients when the physical link  
of the underlying data-link goes down. This needs to be factored-in in  
the logic used to reflect the link status when bridging is configured  
on the underlying data-link.

ngd-03 in bridging-spec.txt 4.1, "However, in the event that bridging  
integrates without Crossbow" since you're not asking for patch binding  
how is this possible?

ngd-04 bridging-spec.txt 4.1 is missing the VLAN VNICs (dladm create- 
vnic -v <vid> ...) in the description.

ngd-05 bridging-design.pdf The design document is still referring to  
old Crossbow architectural details which do not match what was  
integrated in Nevada. For instance, some of the arguments used against  
using the Crossbow classifier don't hold true with the latest Crossbow  
implementation, and the discussion at pages 16-17 should be updated,  
and the design possibly revisited based on the latest Crossbow flow  
table architecture.

Nicolas.

-- 
Nicolas Droux - Solaris Kernel Networking - Sun Microsystems, Inc.
nicolas.droux@sun.com - http://blogs.sun.com/droux


From carlsonj@phorcys.east.sun.com Wed Dec 17 13:00:11 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBHL0AWI000558
	for <psarc-ext@sac.sfbay.Sun.COM>; Wed, 17 Dec 2008 13:00:11 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id mBHL06nk025638;
	Thu, 18 Dec 2008 05:00:08 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KC100919H06JT00@nwk-avmta-1.sfbay.Sun.COM>; Wed,
 17 Dec 2008 13:00:06 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KC1008GHH055B10@nwk-avmta-1.sfbay.Sun.COM>; Wed,
 17 Dec 2008 13:00:06 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id mBHL05Oo008128; Wed,
 17 Dec 2008 16:00:05 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id mBHL05oK008125; Wed,
 17 Dec 2008 16:00:05 -0500 (EST)
Date: Wed, 17 Dec 2008 16:00:05 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: Issues for 2008/055
In-reply-to: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
To: Nicolas Droux <Nicolas.Droux@sun.com>
Cc: PSARC-ext@sun.com
Message-id: <18761.26709.246151.818319@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
Status: RO
Content-Length: 7713

Nicolas Droux writes:
> Here are my issues for Solaris Bridging (PSARC 2008/055).  
> Unfortunately I have a conflict and won't be able to attend the  
> inception review, but I'll be happy to follow-up by email.

I'll integrate these (and my replies) into the existing issues file.

> ngd-01 bridging-spec.txt states "The links assigned to a bridge must  
> not themselves be VLANs, VNICs, or tunnels. Only links that would be  
> acceptable as part of an aggregation or links that are aggregations  
> themselves may be assigned to a bridge." It should be also possible to  
> bridge etherstubs (introduced by Crossbow [PSARC 2006/357]), since  
> they can be used to create virtual switches.

We had an extended talk about this one.  The spec intentionally
doesn't mention etherstubs (except in passing) because they're not
prohibited.

You *should* in principle be able to create a bridge between two
etherstub instances.  I've attempted to do this, and I've found that
there appear to be numerous bugs related to etherstubs in ON today --
for instance, dladm_linkid2legacyname() thinks they're invalid and
dlpi_bind() won't allow me to bind to SAP zero so that I can send and
receive STP into the bit-bucket.

I'm sure I can fix and/or work around those bugs, and thus make it
possible to bridge these objects.  I'll include doing that as part of
the project.  From my prototype:

# dladm show-bridge -l bar
LINK         STATE        UPTIME   DESROOT
stub1        forwarding   22       32768/0:0:0:0:0:0
stub2        forwarding   22       32768/0:0:0:0:0:0

I'm not sure, though, that it's an interesting case.  You'll get
better performance if you just put all of the VNICs that must talk
with each other together on a single etherstub if you're planning to
bridge etherstubs together.  If you're planning to bridge an etherstub
with a regular NIC, then just move the VNICs over to the regular NIC.

> ngd-02 in bridging-spec.txt, 2.2 a), the proposed link/up behavior in  
> the presence of bridges needs to be refined. With Crossbow VNICs, the  
> link status advertised to MAC clients depends also on the presence of  
> other MAC clients on top of the underlying data-link, in order to  
> maintain connectivity between these MAC clients when the physical link  
> of the underlying data-link goes down. This needs to be factored-in in  
> the logic used to reflect the link status when bridging is configured  
> on the underlying data-link.

This appears to be a misunderstanding.  I'm not modifying the existing
link up/down handling that Crossbow VNICs have in any way.

The existing behavior is that the VNIC stays up if there are other
VNICs configured on the same NIC.  The same is true when bridging is
present in the picture: if all of the physical NICs go down, then
VNICs will still do the same thing they did before, and will still
advertise "up" status to clients when there are multiple VNICs present
on the same NIC.

There's no issue here.

> ngd-03 in bridging-spec.txt 4.1, "However, in the event that bridging  
> integrates without Crossbow" since you're not asking for patch binding  
> how is this possible?

The spec was written long ago, before bridging was made dependent on
Crossbow.  I just hadn't removed all the references.

> ngd-04 bridging-spec.txt 4.1 is missing the VLAN VNICs (dladm create- 
> vnic -v <vid> ...) in the description.

I can add that, but the point is the same.  All intentionally
(administratively) created VLANs are the same from the point of view
of this bridging design: they provide (through libdladm) a set of
"allowed VLANs" for the purpose of bridging behavior per 802.1q.

All "casual" VLANs (PPA hack) are different; the bridging code would
use these instances with forwarding disabled for that VLAN.  The user
would have to configure explicitly in order to get forwarding among
the other interfaces.  (That is, PPA hack VLANs do not enter the
"allowed VLAN" set.)

Obviously, since the PPA hack is now gone, the point is moot.  There's
only one kind of VLAN -- the intentionally created kind -- and it
always enters the "allowed VLAN" set.

> ngd-05 bridging-design.pdf The design document is still referring to  
> old Crossbow architectural details which do not match what was  
> integrated in Nevada. For instance, some of the arguments used against  
> using the Crossbow classifier don't hold true with the latest Crossbow  
> implementation, and the discussion at pages 16-17 should be updated,  
> and the design possibly revisited based on the latest Crossbow flow  
> table architecture.

The design document, as I've tried to make clear, is a very early
draft, has not been updated, and is informative for the architectural
review, not normative.  It's explicitly not under review here.

However, the classifier issues still remain, and we discussed those at
length.  The analogy from before still stands: for the same basic
reasons that the Fireengine conn_t classifier can't really be used
effectively as a substitute for the Patricia-tree based IP forwarding
look-up (and vice-versa), the local delivery related classification in
Crossbow doesn't appear suitable for the bridge forwarding case.

With Crossbow, the classification is tied to the administrative bits,
which rely on explicit configuration of the VNICs and flows involved
using a user-space component.  With bridging, forwarding entries are
created and updated on the fly based on source MAC addresses seen in
the data path, and then aged away over time; there's no administrative
involvement normally expected for these entries.

The two are different in many respects.  In theory, though, it might
be possible modify Crossbow so that it can create and destroy
classification entries on the fly (this does not look trivial in the
least; the locking scheme makes this an unobvious approach), and it
may be possible to make use of some aspects of flow administration
when tied to more easily identifiable objects, such as VLANs, though
it's unclear how this should work with the existing Crossbow resource
management structure.

I regard all of that as a research project.  It may well be an
interesting one, but it's not this project by any stretch.  I have no
plans or engineering resources available to redesign the internals of
Crossbow to handle things it wasn't originally designed to do, and I
think that insisting on such an extension of the project I've proposed
is not reasonable.  I will not be doing that.

For what it's worth, it may also be possible to modify Crossbow so
that it eliminates the Fireengine classifier entirely.  After all, the
two are much more aligned than are Crossbow and bridging: both involve
identifying specific receiving client(s) on input and handling output
from multiple clients, and both involve classification structures that
are created strictly on the action of user space components.  It seems
like a performance loss to have Crossbow inspect and classify the
packet once -- potentially looking high up the stack for flow
information -- only to have Fireengine do the same thing again.

I can see that this path wasn't taken, so I can't help but wonder how
reuse of Crossbow's classifier could be considered a requirement for
bridging.

One important issue did come up here: we need to define the relative
ordering between L2 filtering and bridging, and I believe it makes
sense to put L2 filtering closer to the physical I/O.  In other words,
L2 filter should do its work underneath the bridge.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From carlsonj@phorcys.east.sun.com Wed Dec 17 13:45:39 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBHLjc29002308
	for <psarc-ext@sac.sfbay.Sun.COM>; Wed, 17 Dec 2008 13:45:39 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id mBHLjVJd017615;
	Thu, 18 Dec 2008 05:45:35 +0800 (SGT)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KC10050LJ3XPG00@brm-avmta-1.central.sun.com>; Wed,
 17 Dec 2008 14:45:33 -0700 (MST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KC100DF6J3X4JD0@brm-avmta-1.central.sun.com>; Wed,
 17 Dec 2008 14:45:33 -0700 (MST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id mBHLjXjM008544; Wed,
 17 Dec 2008 16:45:33 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id mBHLjXnY008541; Wed,
 17 Dec 2008 16:45:33 -0500 (EST)
Date: Wed, 17 Dec 2008 16:45:33 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: request for comment: 2008/055 Solaris Bridging
To: Darren.Reed@sun.com
Cc: PSARC-ext@sun.com
Message-id: <18761.29437.63549.451206@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
Status: RO
Content-Length: 7909

You had several issues recorded in the 'issues' file for this case,
but you weren't present for the inception review to discuss them.  I
read through the issues and responded as best I could to each, and we
discussed them at some length during the meeting.

I'd like to close this loop to make sure you've had a chance to read
the answers.  Please look over the issues (with written responses)
below, and follow up on any that may not be completely answered.  (You
might also want to visit the minutes and recording of the inception
review itself; some of the verbal answers went beyond the written
ones.  For instance, I made a point of saying that I'm intentionally
not changing anything in the filtering code and that section 6.8 of
the spec is just advice to other project teams.  Since you've read it,
my job is done here.  ;-})

djr-01	From bridging-security.txt (1)(a), it would seem that there is
	potential for an attacker to supply packets to the network that
	would result in excessive CPU use - is that accurate?
	How would an administrator detect this style of problem (melting
	of the network) using Open/Solaris? (Are the observability tools
	provided sufficient?)

Reply:	There are two separate CPU-use-related threats here.  One is
	that an attacker could just flood the network with lots of STP
	traffic for us to handle.  You'd be able to detect that using
	'prstat' and other tools, the effort required to mount the
	attack would be significant, and the impact should be minor.

	The other (which I think you're actually referring to) is the
	network-melting effect of an L2 forwarding loop, which can
	happen if someone can _prevent_ STP packets from getting
	through.

	That's an existing hazard that we all live with in standard
	bridges today: bridged networks sometimes go down because of
	persistent loops, which is one of the reasons why RBridges
	will be better (though that's the subject of a future
	project).

	The change for this project is that instead of being just a
	victim (when other bridges fail), we could be one of the
	active participants in the loop.  If that were the case, you'd
	see the "dladm show-bridge -sl" counters increasing rapidly.

	An inherent problem with this case is that there's no really
	effective way of detecting or countering this failure mode in
	any automatic fashion.  It's not so different from a "really
	busy day."  (Obviously, if detected by a human as an abnormal
	case, shutting down the affected links or bridges will solve
	the problem, and that's usually how it's handled in real
	networks today.)

	The same problem can be caused by the "fastroute $IF" "to $IF"
	and "dup-to $IF" options in ipf.conf when used on the input
	side of a link.  The packet is forwarded to another interface
	without a TTL decrement, and if there's an L2 forwarding path
	between those two interfaces, *exactly* the same failure mode
	occurs.  The difference is that bridging includes Spanning
	Tree, which is designed to detect and disable such loops, and
	IP filter has no such protection.

	It might be possible to advance the state of the art here by
	detecting the combination of high packet rate and identical
	packets (having FCS delivery from the hardware would likely be
	pretty important), but we're not proposing that with this
	project.

djr-02	Given djr-01 and the integration of crossbow to provide MAC layer
	classification and resource controls, is it possible to leverage
	crossbow to protect the system from abuse refered to in (1)(a)?
	If not immediately, is there scope for this as a future project?

Reply:	Crossbow currently identifies flows in MAC clients, such as
	VNICs.  It doesn't work down at the IEEE 802.1 level where
	bridging takes place.

	In principle, it might be possible to create non-flow-oriented
	resource controls down at lower levels, but I don't believe
	that would be a viable fix for the general problem in (1)(a).

	That problem isn't a matter of "abuse" or any bad action on
	the part of other network nodes or users.  It's a matter of a
	single packet being transmitted on a network, picked up by a
	bridge, and then retransmitted on the same network.  Over and
	over again.

	It's not an abuse of the network by someone sending too much
	traffic, so there's place where we can apply a throttle.  It's
	a failure of the network control protocols that ends up making
	the network fundamentally unusable.

	A resource control here would (in theory) limit the rate at
	which we make this forwarding mistake in the case where
	there's a persistent loop, but it wouldn't alleviate the
	problem because all traffic in the same resource class would
	be affected just as though the whole network were swamped.  We
	would still use all of our resource allotment resending the
	same packet over and over.

	For the same reason that you wouldn't ordinarily (at least)
	use a resource control scheme to protect users against
	erroneous IP routes, the same doesn't seem to apply here.

	(If the resource controls here included WFQ among per MAC
	destination queues, then there might be a good argument for
	using that solution.  I think having WFQ with very large
	numbers of output queues would be a good addition to the
	system, regardless of whether bridging is used or not.  It
	doesn't fix the original problem, but it does potentially
	reduce the impact of many classes of DoS problems -- and many
	others as well, such as basic fairness issues in cascaded IP
	forwarding elements.)

djr-03	From bridge-spec.txt, (2.1), the requirement to use individual
	network links to observe packets being sent does not fit with
	what I would expect as a user. Needing to sniff the individual
	network connections seems somewhat onerous (a snoop per link
	in the bridge is required) and presupposes that the "user" knows
	which interface they need to look on for the packet(s) they're
	trying to observe.

Reply:	You can snoop either individual links (if you want to see
	what's going on with that link) or using the special bridge
	observability node described in the section you reference.
	The latter provides a copy of *all* traffic transiting the
	bridge and doesn't require you to snoop individual links.  You
	see everything.

	On Solaris today, you already *do* have to pick a link on
	which you want to snoop, so there's no change in that respect.
	We're adding observability, not taking any away.

djr-05	From bridge-spec.txt, (2.2)(b), how does this impact IP if IP
	interfaces are plumb'd on top of a NIC that is also configured
	to be part of a bridge?

Reply:	IP traffic potentially runs more slowly because it's unable to
	take advantage of hardware features.

	That's a natural consequence of bridging, because the IP stack
	doesn't actually know which underlying link will be used to
	transmit a given packet: that knowledge depends on the bridge
	forwarding table.

	In principle, a future project could handle per-MAC-
	destination hardware options from within IP, but that's not
	this project.  (And probably should be dependent on Erik's
	refactoring project.)

djr-06	From bridge-spec.txt, (2.2)(b), do the notes here apply equally
	to receiving and sending or only to one of the two?

Reply:	They apply to all traffic on the link.

djr-07	From bridge-spec.txt, (6.8), surely you mean that the system
	"should not do this" only when it is being used operationally
	as a bridge.

Reply:	Yes ... though given the possibility of third-party bridging
	code, it seems unwise to allow the case to occur in general
	without at least significant warnings.  Filtering away
	bridging PDUs (STP) looks like a very bad idea to me, and I
	find it hard to see a case where it'd be justified.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Nicolas.Droux@sun.com Wed Dec 17 14:59:38 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBHMxbTG006699
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 17 Dec 2008 14:59:38 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id mBHMxIRY023198
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Wed, 17 Dec 2008 22:59:36 GMT
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KC100B07MJ9Z800@nwk-avmta-2.sfbay.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Wed, 17 Dec 2008 14:59:33 -0800 (PST)
Received: from brmea-mail-1.sun.com ([192.18.98.31])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KC1008KHMJ94A50@nwk-avmta-2.sfbay.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Wed,
 17 Dec 2008 14:59:33 -0800 (PST)
Received: from fe-amer-10.sun.com ([192.18.109.80])
	by brmea-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id mBHMxW0v024793	for
 <PSARC-ext@sun.com>; Wed, 17 Dec 2008 22:59:32 +0000 (GMT)
Received: from conversion-daemon.mail-amer.sun.com by mail-amer.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KC100901JSJBL00@mail-amer.sun.com>
 (original mail from Nicolas.Droux@Sun.COM)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Wed,
 17 Dec 2008 15:59:32 -0700 (MST)
Received: from [10.0.0.2] ([129.150.18.107])
 by mail-amer.sun.com (Sun Java System Messaging Server 6.2-8.04 (built Feb 28
 2007)) with ESMTPSA id <0KC100AYDMIIN790@mail-amer.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Wed,
 17 Dec 2008 15:59:08 -0700 (MST)
Date: Wed, 17 Dec 2008 15:59:01 -0700
From: Nicolas Droux <Nicolas.Droux@sun.com>
Subject: Re: Issues for 2008/055
In-reply-to: <18761.26709.246151.818319@gargle.gargle.HOWL>
Sender: Nicolas.Droux@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: PSARC-ext@sun.com
Message-id: <FEF6920A-4380-46B8-96C1-2D9E471C5C12@Sun.COM>
MIME-version: 1.0
X-Mailer: Apple Mail (2.930.3)
Content-type: text/plain; delsp=yes; format=flowed; charset=US-ASCII
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
 <18761.26709.246151.818319@gargle.gargle.HOWL>
Status: RO
Content-Length: 11015


On Dec 17, 2008, at 2:00 PM, James Carlson wrote:

> Nicolas Droux writes:
>> Here are my issues for Solaris Bridging (PSARC 2008/055).
>> Unfortunately I have a conflict and won't be able to attend the
>> inception review, but I'll be happy to follow-up by email.
>
> I'll integrate these (and my replies) into the existing issues file.

Thanks for your replies. Some follow-up below...

>
>
>> ngd-01 bridging-spec.txt states "The links assigned to a bridge must
>> not themselves be VLANs, VNICs, or tunnels. Only links that would be
>> acceptable as part of an aggregation or links that are aggregations
>> themselves may be assigned to a bridge." It should be also possible  
>> to
>> bridge etherstubs (introduced by Crossbow [PSARC 2006/357]), since
>> they can be used to create virtual switches.
>
> We had an extended talk about this one.  The spec intentionally
> doesn't mention etherstubs (except in passing) because they're not
> prohibited.
>
> You *should* in principle be able to create a bridge between two
> etherstub instances.  I've attempted to do this, and I've found that
> there appear to be numerous bugs related to etherstubs in ON today --
> for instance, dladm_linkid2legacyname() thinks they're invalid and
> dlpi_bind() won't allow me to bind to SAP zero so that I can send and
> receive STP into the bit-bucket.

I don't think the dladm_linkid2legacyname() you are seeing is a bug.  
As its name implies, it is used for legacy data-link names and  
etherstubs don't fall in that category.

We also limitations built-in to prevent an etherstub to be plumbed,  
which maybe causing the other issue you are hitting.

So things seem to be currently working as expected. If there are new  
requirements for etherstubs in order to make them work with bridging,  
we'll be happy to work with you on that.

> I'm sure I can fix and/or work around those bugs, and thus make it
> possible to bridge these objects.  I'll include doing that as part of
> the project.  From my prototype:
>
> # dladm show-bridge -l bar
> LINK         STATE        UPTIME   DESROOT
> stub1        forwarding   22       32768/0:0:0:0:0:0
> stub2        forwarding   22       32768/0:0:0:0:0:0
>
> I'm not sure, though, that it's an interesting case.  You'll get
> better performance if you just put all of the VNICs that must talk
> with each other together on a single etherstub if you're planning to
> bridge etherstubs together.  If you're planning to bridge an etherstub
> with a regular NIC, then just move the VNICs over to the regular NIC.

An important benefit is to have the flexibility to build virtual  
networks in a box which map directly to physical topologies.


>> ngd-02 in bridging-spec.txt, 2.2 a), the proposed link/up behavior in
>> the presence of bridges needs to be refined. With Crossbow VNICs, the
>> link status advertised to MAC clients depends also on the presence of
>> other MAC clients on top of the underlying data-link, in order to
>> maintain connectivity between these MAC clients when the physical  
>> link
>> of the underlying data-link goes down. This needs to be factored-in  
>> in
>> the logic used to reflect the link status when bridging is configured
>> on the underlying data-link.
>
> This appears to be a misunderstanding.  I'm not modifying the existing
> link up/down handling that Crossbow VNICs have in any way.

I think it's the following sentence in your document which which is  
confusing to me: "This means that when all external links are showing  
link-down status, the upper-level clients using the MAC layers will  
see link-down events as well."

VNICs are MAC clients, and their link status may not reflect the link- 
down events of the external links.

> The existing behavior is that the VNIC stays up if there are other
> VNICs configured on the same NIC.  The same is true when bridging is
> present in the picture: if all of the physical NICs go down, then
> VNICs will still do the same thing they did before, and will still
> advertise "up" status to clients when there are multiple VNICs present
> on the same NIC.
>
> There's no issue here.

I'd suggest clarifying the interactions in the spec.

>> ngd-05 bridging-design.pdf The design document is still referring to
>> old Crossbow architectural details which do not match what was
>> integrated in Nevada. For instance, some of the arguments used  
>> against
>> using the Crossbow classifier don't hold true with the latest  
>> Crossbow
>> implementation, and the discussion at pages 16-17 should be updated,
>> and the design possibly revisited based on the latest Crossbow flow
>> table architecture.
>
> The design document, as I've tried to make clear, is a very early
> draft, has not been updated, and is informative for the architectural
> review, not normative.  It's explicitly not under review here.
>
> However, the classifier issues still remain, and we discussed those at
> length.  The analogy from before still stands: for the same basic
> reasons that the Fireengine conn_t classifier can't really be used
> effectively as a substitute for the Patricia-tree based IP forwarding
> look-up (and vice-versa), the local delivery related classification in
> Crossbow doesn't appear suitable for the bridge forwarding case.
>
> With Crossbow, the classification is tied to the administrative bits,
> which rely on explicit configuration of the VNICs and flows involved
> using a user-space component.  With bridging, forwarding entries are
> created and updated on the fly based on source MAC addresses seen in
> the data path, and then aged away over time; there's no administrative
> involvement normally expected for these entries.

This doesn't have to be the case. The Crossbow flow implementation  
provides a kernel API which allows flows to be created and added to  
flow tables. That API today is used for VNICs, but also via MAC client  
creation in general (e.g. through LDOMs), for user-specified flows,  
and for multicast addresses. It could be used by the bridge code as  
well.

> The two are different in many respects.  In theory, though, it might
> be possible modify Crossbow so that it can create and destroy
> classification entries on the fly (this does not look trivial in the
> least; the locking scheme makes this an unobvious approach), and it
> may be possible to make use of some aspects of flow administration
> when tied to more easily identifiable objects, such as VLANs, though
> it's unclear how this should work with the existing Crossbow resource
> management structure.

Crossbow can already create and destroy flow entries on the fly. The  
locking requirements are also very straightforward.

> I regard all of that as a research project.  It may well be an
> interesting one, but it's not this project by any stretch.  I have no
> plans or engineering resources available to redesign the internals of
> Crossbow to handle things it wasn't originally designed to do, and I
> think that insisting on such an extension of the project I've proposed
> is not reasonable.  I will not be doing that.

I don't think you need to "redesign the internals of Crossbow".

We have kernel APIs which I believe can achieve most of what you need  
here. There might be some small gaps, but I don't see why you would  
need to introduce a new classification table at layer 2 since we  
already have most of what you need at the same layer in mac.

> For what it's worth, it may also be possible to modify Crossbow so
> that it eliminates the Fireengine classifier entirely.  After all, the
> two are much more aligned than are Crossbow and bridging: both involve
> identifying specific receiving client(s) on input and handling output
> from multiple clients, and both involve classification structures that
> are created strictly on the action of user space components.  It seems
> like a performance loss to have Crossbow inspect and classify the
> packet once -- potentially looking high up the stack for flow
> information -- only to have Fireengine do the same thing again.

Of course it might be possible to use flows from other layers of the  
stack, but this is not as obvious as bridging. See below...

> I can see that this path wasn't taken, so I can't help but wonder how
> reuse of Crossbow's classifier could be considered a requirement for
> bridging.

It is very relevant to bridging since the bridge forwarding happens at  
the same place on the data-path as the classification that Crossbow  
introduced in the MAC layer. For example on transmit, the  
classification on the destination MAC address results in sending the  
packet to another MAC client (e.g. VNIC), send copies of the packets  
to members of a multicast group, or send the packet on the wire. A new  
outcome would be to pass the packet to a bridge.

With Crossbow the old mac txinfo implementation is completely gone,  
and all packets sent from a client will go through mac_tx(). Since  
mac_tx() is where classification takes place, and where you need to do  
your own checks, it seems natural to combine the operations in a  
single classification operation.

Similarly on receive, the old mac rx_add() entry points are gone, and  
demultiplexing to the interested parties is now done by the mac layer  
through the same classification table. So having the entry for an  
address the bridge is interested in would allow the classification to  
be leveraged for the receive side as well.

So reusing the classifier for components of the same layer of the  
stack seems that the natural thing to do. Once you use flows you can  
also take advantage of hardware classification on the receive side.

The Crossbow team will be happy to answer questions you may have about  
the new datapath, and discuss specific requirements you may have, and  
are not addressed by the current implementation.

> One important issue did come up here: we need to define the relative
> ordering between L2 filtering and bridging, and I believe it makes
> sense to put L2 filtering closer to the physical I/O.  In other words,
> L2 filter should do its work underneath the bridge.

There's filtering which needs to occur between multiple MAC clients  
(VNICs are MAC clients) defined on top of the same data-link. For  
example to be consistent with the way things work in the physical  
world, one might want to prevent a VM to be able to specific send  
packets on the wire, which in this case would include a bridge. On the  
transmit side these checks would have to be done before the packet is  
potentially sent through a bridge, i.e. the L2 filtering would have to  
be done "on top" of the bridge.

Nicolas.

>
>
> -- 
> James Carlson, Solaris Networking              <james.d.carlson@sun.com 
> >
> Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442  
> 2084
> MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442  
> 1677

-- 
Nicolas Droux - Solaris Kernel Networking - Sun Microsystems, Inc.
nicolas.droux@sun.com - http://blogs.sun.com/droux


From Darren.Reed@sun.com Wed Dec 17 19:30:07 2008
Received: from sunmail2sca.sfbay.sun.com (sunmail2sca.SFBay.Sun.COM [129.145.155.234])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBI3U7DS029423
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 17 Dec 2008 19:30:07 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail2sca.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mBI3U2W0022046
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Wed, 17 Dec 2008 19:30:07 -0800 (PST)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KC10072XZ263N00@nwk-avmta-2.sfbay.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Wed, 17 Dec 2008 19:30:06 -0800 (PST)
Received: from gmp-eb-inf-2.sun.com ([192.18.6.24])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KC100J6LZ25R960@nwk-avmta-2.sfbay.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Wed,
 17 Dec 2008 19:30:06 -0800 (PST)
Received: from fe-emea-09.sun.com (gmp-eb-lb-1-fe3.eu.sun.com [192.18.6.10])
	by gmp-eb-inf-2.sun.com (8.13.7+Sun/8.12.9) with ESMTP id mBI3U5mu029383	for
 <PSARC-ext@sun.com>; Thu, 18 Dec 2008 03:30:05 +0000 (GMT)
Received: from conversion-daemon.fe-emea-09.sun.com by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KC100001YXVF000@fe-emea-09.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Thu,
 18 Dec 2008 03:30:05 +0000 (GMT)
Received: from [129.158.90.191] by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KC1002JXZ22LW70@fe-emea-09.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Thu, 18 Dec 2008 03:30:05 +0000 (GMT)
Date: Thu, 18 Dec 2008 14:29:54 +1100
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: request for comment: 2008/055 Solaris Bridging
In-reply-to: <18761.29437.63549.451206@gargle.gargle.HOWL>
Sender: Darren.Reed@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: PSARC-ext@sun.com
Message-id: <4949C3B2.7020405@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <18761.29437.63549.451206@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.18 (Windows/20081105)
Status: RO
Content-Length: 2977

James Carlson wrote:
> ...
> djr-02	Given djr-01 and the integration of crossbow to provide MAC layer
> 	classification and resource controls, is it possible to leverage
> 	crossbow to protect the system from abuse refered to in (1)(a)?
> 	If not immediately, is there scope for this as a future project?
>
> Reply:	Crossbow currently identifies flows in MAC clients, such as
> 	VNICs.  It doesn't work down at the IEEE 802.1 level where
> 	bridging takes place.

So it isn't possible to use Crossbow's interfaces to
put STP packets into a separate rx/tx ring pair or to
otherwise use crossbow to partition rx/tx rings up for
preferential treatment of specific ethernet addresses
on either side of a bridge? And thus if we can do that,
then it seems to me like we should be able to specify
what sort of bandwidth allocation/guarantees those
rings get...

Or is this RFE material?


> djr-03	From bridge-spec.txt, (2.1), the requirement to use individual
> 	network links to observe packets being sent does not fit with
> 	what I would expect as a user. Needing to sniff the individual
> 	network connections seems somewhat onerous (a snoop per link
> 	in the bridge is required) and presupposes that the "user" knows
> 	which interface they need to look on for the packet(s) they're
> 	trying to observe.
>
> Reply:	You can snoop either individual links (if you want to see
> 	what's going on with that link) or using the special bridge
> 	observability node described in the section you reference.
> 	The latter provides a copy of *all* traffic transiting the
> 	bridge and doesn't require you to snoop individual links.  You
> 	see everything.
>
> 	On Solaris today, you already *do* have to pick a link on
> 	which you want to snoop, so there's no change in that respect.
> 	We're adding observability, not taking any away.

The distinction I'm keen to make is observing received packets
vs sent packets. This is the paragraph that I'm referring to:

"To see the packets transmitted and received on a particular link
 (after the bridging process is complete), snoop on the individual
 links rather than the bridge observability node."

What I'm not sure about is whether "handled by the bridge" in the
other paragraphs in this section refers to packets that are both
sent and received, just received, or something else. This, in
concert with promiscuous mode being required with snoop to get
sent packets with DLPI, has me asking for this to be more clear,
especially considering this sentence:

"The packets delivered will represent the data received by the bridge."

I think this section needs to make it clear whether snoop or the
observability devices will present:
1) traffic that is received by the bridge
2) traffic that is transmitted by the bridge (both STP + data)
3) traffic that is accepted/forwarded by the bridge

i.e. if I'm snoop'ing bridge0 and a packet comes in bge0 and
the bridge sends it out bge1, will I see it once with snoop or
twice or...?

Darren


From carlsonj@phorcys.east.sun.com Thu Dec 18 08:25:58 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBIGPv53018256
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 18 Dec 2008 08:25:58 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id mBIGPrNj021733;
	Thu, 18 Dec 2008 16:25:55 GMT
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KC20030BYZ5OE00@nwk-avmta-1.sfbay.Sun.COM>; Thu,
 18 Dec 2008 08:25:53 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KC2009SKYZ45ZD0@nwk-avmta-1.sfbay.Sun.COM>; Thu,
 18 Dec 2008 08:25:53 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id mBIGPqS4007065; Thu,
 18 Dec 2008 11:25:52 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id mBIGPqLt007062; Thu,
 18 Dec 2008 11:25:52 -0500 (EST)
Date: Thu, 18 Dec 2008 11:25:52 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: Issues for 2008/055
In-reply-to: <FEF6920A-4380-46B8-96C1-2D9E471C5C12@Sun.COM>
To: Nicolas Droux <Nicolas.Droux@sun.com>
Cc: PSARC-ext@sun.com
Message-id: <18762.31120.53165.970797@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
 <18761.26709.246151.818319@gargle.gargle.HOWL>
 <FEF6920A-4380-46B8-96C1-2D9E471C5C12@Sun.COM>
Status: RO
Content-Length: 18889

Nicolas Droux writes:
> > You *should* in principle be able to create a bridge between two
> > etherstub instances.  I've attempted to do this, and I've found that
> > there appear to be numerous bugs related to etherstubs in ON today --
> > for instance, dladm_linkid2legacyname() thinks they're invalid and
> > dlpi_bind() won't allow me to bind to SAP zero so that I can send and
> > receive STP into the bit-bucket.
> 
> I don't think the dladm_linkid2legacyname() you are seeing is a bug.  
> As its name implies, it is used for legacy data-link names and  
> etherstubs don't fall in that category.

I see ... dladm_datalink_id2info seems like the right way out of this
one.

> We also limitations built-in to prevent an etherstub to be plumbed,  
> which maybe causing the other issue you are hitting.

That seems likely.

The underlying issue seems to be in the design of VNICs and etherstub.
On a regular NIC, there's a native instance that has its own MAC
address and thus can work normally when plumbed; no VNIC is required.
On an etherstub, though, the native instance has no address, and VNICs
appear to be required in order to provide MAC addresses needed for
clients.

Thus, to a client, the two are different.

Providing etherstubs themselves with arbitrary MAC addresses and
allowing them to be plumbed normally (just as though they were regular
NICs) would resolve the issue, but might be a step too far.

> So things seem to be currently working as expected. If there are new  
> requirements for etherstubs in order to make them work with bridging,  
> we'll be happy to work with you on that.

Bridging works with 802 MAC entities and the things that emulate them,
such as aggregations.  As long as an "etherstub" can be made to look
like a real Ethernet (with transmit just going to the bit-bucket and
nothing ever received from "the wire"), it should work fine.

Currently, it doesn't seem to look much like an Ethernet by itself.

> > I'm not sure, though, that it's an interesting case.  You'll get
> > better performance if you just put all of the VNICs that must talk
> > with each other together on a single etherstub if you're planning to
> > bridge etherstubs together.  If you're planning to bridge an etherstub
> > with a regular NIC, then just move the VNICs over to the regular NIC.
> 
> An important benefit is to have the flexibility to build virtual  
> networks in a box which map directly to physical topologies.

Please expand on that.

If we're talking about using these for general testing purposes (e.g.,
creating a four-port bridge inside a box and using that for the IPMP
test suite; no external hardware required), then I think etherstubs
are just the wrong mechanism, because they are so different from
actual Ethernets.

What I believe we really need to generalize that case is an entity
that emulates a point-to-point link, where the two endpoints are
separate MACs in the same box.  Such a pseudo-device would give us:

  - Distinct transmit and receive paths across the emulated link,
    which could be subject to any desired impairment (e.g., MTU
    restrictions, loss rates, delay, link up/down).

  - Emulation of negotiated L2 parameters, including all of the
    standard Ethernet capabilities.

  - An obvious means to test wireless configurations, by having one
    endpoint behave as an 802.11 "client" and the other behave as an
    AP, emulating the assocation/disassociation/authentication
    mechanisms.

None of that appears to be as feasible with etherstubs.

If we're talking about other cases, then I'll need more information.
One case suggested by private email was a DomU migrating from one NIC
to another, and using an etherstub bridged to the actual NIC in order
to make the transition "easy."  But if that's the goal, I would
suggest that having either a mechanism that allows migration of VNICs
among NICs without teardown or something that inserts an indirection
mechanism (perhaps even an etherstub with the optional capability of
linking to an underlying NIC) would be better.

Don't forget that NICs used in bridges _must_ be in promiscuous mode,
and that this destroys a good bit of performance.  That alone makes
them less interesting for ad-hoc system reconfiguration types of
activities.

> > This appears to be a misunderstanding.  I'm not modifying the existing
> > link up/down handling that Crossbow VNICs have in any way.
> 
> I think it's the following sentence in your document which which is  
> confusing to me: "This means that when all external links are showing  
> link-down status, the upper-level clients using the MAC layers will  
> see link-down events as well."

Will fix ... the "upper-level" here means everything that thinks it's
looking directly at the drivers.  MAC client and VNIC behavior doesn't
change.

> > With Crossbow, the classification is tied to the administrative bits,
> > which rely on explicit configuration of the VNICs and flows involved
> > using a user-space component.  With bridging, forwarding entries are
> > created and updated on the fly based on source MAC addresses seen in
> > the data path, and then aged away over time; there's no administrative
> > involvement normally expected for these entries.
> 
> This doesn't have to be the case. The Crossbow flow implementation  
> provides a kernel API which allows flows to be created and added to  
> flow tables. That API today is used for VNICs, but also via MAC client  
> creation in general (e.g. through LDOMs), for user-specified flows,  
> and for multicast addresses. It could be used by the bridge code as  
> well.

The mac_flow_add() function asserts that the ft_mip perimeter is
held.  My understanding of how the perimeters are used in Crossbow
(please correct me if I've gotten this wrong) is that they're *NEVER*
taken in the upward direction, as that would lead to deadlock.  This
is given as general rule "R2" in the mac.c block comments.

Assuming that to be correct, it would mean that potentially _all_
received packets would have to be shuffled off to a separate kernel
thread, so that the source MAC address could be safely inspected and
used to update the list of learned MAC destinations.

Worse still, it depends on a per-MAC lock for a per-MAC flow table,
which makes no sense for bridging.  Bridge forwarding entries are a
common database for a given bridge instance -- packets flowing in or
out of any port on the bridge are matched against the common pool, not
against a set of per-port entries.  We thus need a resource common to
multiple MACs, which is not something that appears to be part of
Crossbow, or we need to multiply the storage required by replicating
each entry by N (and the locking effort goes up by N as well), so we
can remain with Crossbow's per-port mi_flow_tab.

The scheme I'm currently using is not so drastic.  I'm not holding the
rw lock I'm using across any external calls (other than the avl
functions), so I'm able to update the table in the datapath when
necessary by taking a writer lock to insert a new entry.  And my
database is an avl tree attached to the bridge instance, so it's
common among the attached ports.

The forwarding entries I have are more like ire_t entries than they
are like conn_ts.  The flows you have look more like conn_ts to me.  I
think they're different objects.

> > The two are different in many respects.  In theory, though, it might
> > be possible modify Crossbow so that it can create and destroy
> > classification entries on the fly (this does not look trivial in the
> > least; the locking scheme makes this an unobvious approach), and it
> > may be possible to make use of some aspects of flow administration
> > when tied to more easily identifiable objects, such as VLANs, though
> > it's unclear how this should work with the existing Crossbow resource
> > management structure.
> 
> Crossbow can already create and destroy flow entries on the fly. The  
> locking requirements are also very straightforward.

They're straightforward but apparently immiscible with the data-driven
learning required for bridging.

> > I regard all of that as a research project.  It may well be an
> > interesting one, but it's not this project by any stretch.  I have no
> > plans or engineering resources available to redesign the internals of
> > Crossbow to handle things it wasn't originally designed to do, and I
> > think that insisting on such an extension of the project I've proposed
> > is not reasonable.  I will not be doing that.
> 
> I don't think you need to "redesign the internals of Crossbow".
> 
> We have kernel APIs which I believe can achieve most of what you need  
> here. There might be some small gaps, but I don't see why you would  
> need to introduce a new classification table at layer 2 since we  
> already have most of what you need at the same layer in mac.

I've stated my needs previously, and I don't see how Crossbow's
features would fill them or be modified to fill them.  To state them
again in detail:

  - I need to be able to add and delete entries in the table from
    within the datapath; this means manipulating the flow table on an
    upcall from the driver layer.  This means having a locking scheme
    that supports table modification in an upward direction.

  - I need to control the forwarding function on a per-port basis in
    order to implement Spanning Tree.  When disabled (the default),
    arriving packets are never matched against bridge forwarding
    entries, and are just delivered to local destinations (i.e., VNICs
    on that port).  When forwarding is enabled by STP, arriving
    packets are matched against forwarding entries and delivered as
    directed.  This means we need to match against subsets of flows
    (local to port or all flows) depending on port state.

  - I need to do a special match on "unknown destinations" -- anything
    not in the forwarding table must be copied to every port that's
    enabled by the control described above.

  - The forwarding table must be implemented in a per-bridge manner,
    rather than per-port/link, as is currently done with Crossbow.
    Forwarding is a global function for a bridge; otherwise, you can't
    actually forward between ports.

  - If I'm to support IVL and SVL, I need to vary my lookup so that it
    sometimes matches on MAC address alone (SVL) and other times based
    on MAC+VID (IVL).  (I could ditch the feature and hard code for
    one or the other, but that'd be less good.)

  - I need to age away entries (potentially based on bridge
    parameters).  (This part at least looks "easy" in that I would use
    a separate kernel thread that wakes up when needed.)

There are many other issues that I don't know how to resolve.  For a
few of these:

  - Crossbow flows are associated with resources (CPUs and the like),
    and these are currently assigned when flows are created by
    administrative action, but it's unclear what choices data-driven
    flow creation should use for those parameters.  There's probably a
    separate infrastructure needed here -- perhaps per-VLAN set of
    flow 'templates' set up by an administrator that get used to
    create the actual flows -- but it's unclear how that should work.

  - What happens when L2 forwarding occurs?  Do we use the same input
    side flow for the output side, or are two separate look-ups done?
    Things are clearer for IP and other network layer features using
    Crossbow, as the xb responsibility effectively ends at the network
    layer, so the user naturally expects a new flow look-up if the
    packet is transmitted elsewhere by IP.  It's unclear if that makes
    sense for forwarding done entirely inside Crossbow.

  - Is "unknown destination" (transmit on all ports) a single flow or
    N separate output-side look-ups?

  - What happens when the user configures MAC based flows *and*
    bridging is in use?  We then have two administrators manipulating
    the same table -- the human admin is creating static entries, and
    the system is creating and deleting dynamic ones.  What happens
    when they overlap?

> > For what it's worth, it may also be possible to modify Crossbow so
> > that it eliminates the Fireengine classifier entirely.  After all, the
> > two are much more aligned than are Crossbow and bridging: both involve
> > identifying specific receiving client(s) on input and handling output
> > from multiple clients, and both involve classification structures that
> > are created strictly on the action of user space components.  It seems
> > like a performance loss to have Crossbow inspect and classify the
> > packet once -- potentially looking high up the stack for flow
> > information -- only to have Fireengine do the same thing again.
> 
> Of course it might be possible to use flows from other layers of the  
> stack, but this is not as obvious as bridging. See below...

I have to differ on that.  Flow identification is exactly what
Fireengine aims to provide.  It's also what Crossbow provides.
Matching them together means that you don't have to do the same
look-up twice, and means that you can extend resource controls through
to individual sockets, applications, and users.  I think it
potentially also means that Crossbow can provide squeue-like
functionality on behalf of IP, so that a whole layer of locks and
complexity can be removed, straight from NIC to transport.

The use of Crossbow inside bridging is much less obvious to me.  I
certainly agree that it's possible.  After having looked at bridging
for some time now, the application of Crossbow structures looks to me
like a research project and involves non-trivial redesign of Crossbow.

> > I can see that this path wasn't taken, so I can't help but wonder how
> > reuse of Crossbow's classifier could be considered a requirement for
> > bridging.
> 
> It is very relevant to bridging since the bridge forwarding happens at  
> the same place on the data-path as the classification that Crossbow  
> introduced in the MAC layer. For example on transmit, the  
> classification on the destination MAC address results in sending the  
> packet to another MAC client (e.g. VNIC), send copies of the packets  
> to members of a multicast group, or send the packet on the wire. A new  
> outcome would be to pass the packet to a bridge.

The flow identification in flow_ip_v4_match() does the same sort of
look-up that's done by IPv4 forwarding entries, IP Filter, and IPsec
policy matching.  That doesn't make Crossbow a replacement for any of
those other cases, though.

At a high enough level, the two functions do appear to be similar.
They're not the same, though, and have significant differences when
you look at the details.  For instance, the things I'm matching
against are *not* per-port entries.

> With Crossbow the old mac txinfo implementation is completely gone,  
> and all packets sent from a client will go through mac_tx(). Since  
> mac_tx() is where classification takes place, and where you need to do  
> your own checks, it seems natural to combine the operations in a  
> single classification operation.

Yes, that makes it "possible."  The issues discussed above make it
much less natural.

> Similarly on receive, the old mac rx_add() entry points are gone, and  
> demultiplexing to the interested parties is now done by the mac layer  
> through the same classification table. So having the entry for an  
> address the bridge is interested in would allow the classification to  
> be leveraged for the receive side as well.
> 
> So reusing the classifier for components of the same layer of the  
> stack seems that the natural thing to do. Once you use flows you can  
> also take advantage of hardware classification on the receive side.
> 
> The Crossbow team will be happy to answer questions you may have about  
> the new datapath, and discuss specific requirements you may have, and  
> are not addressed by the current implementation.

Again, I have no plan or resources available to launch the necessary
research project into the redesign of Crossbow for the use of
bridging.  It sounds like an interesting thing, to be sure, but that's
just not going to happen within the context of this project.

If you feel a TCR or outright denial of this project is necessary to
preserve Crossbow, then speak with the other ARC members and build a
consensus for that.

If someone else has the necessary resources to do the required
research (or even wants to launch a counter-project), then be my
guest.

> > One important issue did come up here: we need to define the relative
> > ordering between L2 filtering and bridging, and I believe it makes
> > sense to put L2 filtering closer to the physical I/O.  In other words,
> > L2 filter should do its work underneath the bridge.
> 
> There's filtering which needs to occur between multiple MAC clients  
> (VNICs are MAC clients) defined on top of the same data-link. For  
> example to be consistent with the way things work in the physical  
> world, one might want to prevent a VM to be able to specific send  
> packets on the wire, which in this case would include a bridge. On the  
> transmit side these checks would have to be done before the packet is  
> potentially sent through a bridge, i.e. the L2 filtering would have to  
> be done "on top" of the bridge.

The issue that I'm pointing out is largely an administrative one.
When someone creates an L2 filter, what exactly are they expecting to
filter on?

If they're expecting to filter on the actual physical interface that
bridging uses, then the scheme you're suggesting won't work right --
the action of bridge forwarding can (and does) redirect output packets
from the originally intended link over to the actual one where the
destination resides -- or to all ports if the destination is not
known.  If an administrator says "never send packets to
00:01:02:03:04:05 on hme0", then he'll be quite surprised to find that
when he sends packets to that address on ce0, the packet comes
stumbling out hme0 despite his filter, because the L2 filtering "on
top" of the bridge saw only ce0 as the output.

Perhaps a hybrid approach is needed.  Place the hooks in two places:
below the bridge for physical I/O and inside the MAC layer for
client-to-client (VNIC-to-VNIC) communication.  The former will allow
the user to filter packets with reference to the actual physical
hardware in use (as opposed to just what the network layer "thinks" is
in use), and allow him to filter traffic between MAC clients as well.
The latter would not involve bridging, and thus would use the link as
known to the network layer.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From carlsonj@phorcys.east.sun.com Thu Dec 18 08:51:45 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBIGpimd018551
	for <psarc-ext@sac.sfbay.Sun.COM>; Thu, 18 Dec 2008 08:51:44 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id mBIGpapC023934;
	Fri, 19 Dec 2008 00:51:42 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KC300611065WX00@nwk-avmta-1.sfbay.Sun.COM>; Thu,
 18 Dec 2008 08:51:41 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KC3009510646AF0@nwk-avmta-1.sfbay.Sun.COM>; Thu,
 18 Dec 2008 08:51:41 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id mBIGpex3007153; Thu,
 18 Dec 2008 11:51:40 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id mBIGperZ007150; Thu,
 18 Dec 2008 11:51:40 -0500 (EST)
Date: Thu, 18 Dec 2008 11:51:40 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: request for comment: 2008/055 Solaris Bridging
In-reply-to: <4949C3B2.7020405@Sun.COM>
To: Darren Reed <Darren.Reed@sun.com>
Cc: PSARC-ext@sun.com
Message-id: <18762.32668.469762.480772@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <18761.29437.63549.451206@gargle.gargle.HOWL>
 <4949C3B2.7020405@Sun.COM>
Status: RO
Content-Length: 5279

Darren Reed writes:
> James Carlson wrote:
> > ...
> > djr-02	Given djr-01 and the integration of crossbow to provide MAC layer
> > 	classification and resource controls, is it possible to leverage
> > 	crossbow to protect the system from abuse refered to in (1)(a)?
> > 	If not immediately, is there scope for this as a future project?
> >
> > Reply:	Crossbow currently identifies flows in MAC clients, such as
> > 	VNICs.  It doesn't work down at the IEEE 802.1 level where
> > 	bridging takes place.
> 
> So it isn't possible to use Crossbow's interfaces to
> put STP packets into a separate rx/tx ring pair or to
> otherwise use crossbow to partition rx/tx rings up for
> preferential treatment of specific ethernet addresses
> on either side of a bridge? And thus if we can do that,
> then it seems to me like we should be able to specify
> what sort of bandwidth allocation/guarantees those
> rings get...
> 
> Or is this RFE material?

You could certainly do that if you wanted to provide some Crossbow
protection to the STP traffic itself.  That won't do much to repair
things if an L2 loop is inadvertently introduced, but it might
possibly lessen the likelihood of one happening.

I thought you were talking more about the data path fault than the
control path one.  In the data path, I don't think there's much that
can be done if a live L2 loop exists due to control path failure.

(I'm not sure if it makes sense for us to define static flows or
resources to use "automatically" when STP is present.  Perhaps that's
something that can be included in the documentation instead -- a
section on using Crossbow to protect the control path if desired.)

> > djr-03	From bridge-spec.txt, (2.1), the requirement to use individual
> > 	network links to observe packets being sent does not fit with
> > 	what I would expect as a user. Needing to sniff the individual
> > 	network connections seems somewhat onerous (a snoop per link
> > 	in the bridge is required) and presupposes that the "user" knows
> > 	which interface they need to look on for the packet(s) they're
> > 	trying to observe.
[...]
> The distinction I'm keen to make is observing received packets
> vs sent packets.

There's no real distinction here.

> This is the paragraph that I'm referring to:
> 
> "To see the packets transmitted and received on a particular link
>  (after the bridging process is complete), snoop on the individual
>  links rather than the bridge observability node."
> 
> What I'm not sure about is whether "handled by the bridge" in the
> other paragraphs in this section refers to packets that are both
> sent and received, just received, or something else.

Both.  I'll update this section to make it clearer.

> This, in
> concert with promiscuous mode being required with snoop to get
> sent packets with DLPI, has me asking for this to be more clear,
> especially considering this sentence:
> 
> "The packets delivered will represent the data received by the bridge."

This means "received in any fashion" -- through downward calls
(transmits made by MAC clients) or by upward calls (packets received
by drivers).

> I think this section needs to make it clear whether snoop or the
> observability devices will present:
> 1) traffic that is received by the bridge
> 2) traffic that is transmitted by the bridge (both STP + data)
> 3) traffic that is accepted/forwarded by the bridge
> 
> i.e. if I'm snoop'ing bridge0 and a packet comes in bge0 and
> the bridge sends it out bge1, will I see it once with snoop or
> twice or...?

You should see it once.

If you snoop on the bridge observability node, you see a copy of each
packet that goes through the bridge code.  That includes:

  - All packets received by the driver for any link assigned to the
    bridge, as long as the link itself is in "forwarding" state.
    (Ports disabled by STP are as though they don't exist in terms of
    the data path for the bridge.)

  - All packets transmitted by a MAC client on any link assigned to
    the bridge, as long as the link itself is in "forwarding" state.

  - Since STP is itself just a plain old DLPI user, you will see the
    STP packets on all of the enabled ports when snooping the bridge
    observability node.  (We could filter these away if desired, as
    they're not bridge data path items, but there didn't seem to be a
    need to do this.)

If you snoop on a link, you see what's transmitted or received on that
one link, as far as the MAC and higher layers are concerned, and
without regard to the bridge forwarding state.  Again, as STP is just
a DLPI user, that means you'll see the STP packets for a given port,
even if the port is disabled by STP.  (I need to make that part much
clearer.)

The operation of the observability node on the bridge (with the
possible exception of the STP messages) is intended to be just like a
mirroring port on a regular bridge.  The difference between this and
having a real mirroring port is that it's a virtual interface that's
established dynamically, so it has no overhead when not in use.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From droux@sun.com Mon Dec 22 23:11:19 2008
Received: from sunmail3mpk.sfbay.sun.com (sunmail3mpk.SFBay.Sun.COM [129.146.11.52])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBN7BJr4005226
	for <psarc-ext@sac.sfbay.sun.com>; Mon, 22 Dec 2008 23:11:19 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail3mpk.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mBN7B6GU025355
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Mon, 22 Dec 2008 23:11:19 -0800 (PST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KCB00J13IMLSC00@brm-avmta-1.central.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Tue, 23 Dec 2008 00:11:09 -0700 (MST)
Received: from brmea-mail-1.sun.com ([192.18.98.31])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KCB00C0VIMKVZC0@brm-avmta-1.central.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Tue,
 23 Dec 2008 00:11:08 -0700 (MST)
Received: from fe-amer-09.sun.com ([192.18.109.79])
	by brmea-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id mBN7B8k5013691	for
 <PSARC-ext@sun.com>; Tue, 23 Dec 2008 07:11:08 +0000 (GMT)
Received: from conversion-daemon.mail-amer.sun.com by mail-amer.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KCB00701IHLDZ00@mail-amer.sun.com> (original mail from droux@sun.com)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Tue,
 23 Dec 2008 00:11:07 -0700 (MST)
Received: from [10.0.0.2] ([97.119.179.69])
 by mail-amer.sun.com (Sun Java System Messaging Server 6.2-8.04 (built Feb 28
 2007)) with ESMTPSA id <0KCB005HQIMI6TE0@mail-amer.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Tue,
 23 Dec 2008 00:11:07 -0700 (MST)
Date: Tue, 23 Dec 2008 00:11:01 -0700
From: Nicolas Droux <droux@sun.com>
Subject: Re: Issues for 2008/055
In-reply-to: <18762.31120.53165.970797@gargle.gargle.HOWL>
Sender: Nicolas.Droux@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: Nicolas Droux <Nicolas.Droux@sun.com>, PSARC-ext@sun.com
Message-id: <1491995A-D70F-4856-9627-36EC11B0B8CF@sun.com>
MIME-version: 1.0
X-Mailer: Apple Mail (2.930.3)
Content-type: text/plain; delsp=yes; format=flowed; charset=US-ASCII
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
 <18761.26709.246151.818319@gargle.gargle.HOWL>
 <FEF6920A-4380-46B8-96C1-2D9E471C5C12@Sun.COM>
 <18762.31120.53165.970797@gargle.gargle.HOWL>
Status: RO
Content-Length: 19286


On Dec 18, 2008, at 9:25 AM, James Carlson wrote:

> Nicolas Droux writes:
>>
>> We also limitations built-in to prevent an etherstub to be plumbed,
>> which maybe causing the other issue you are hitting.
>
> That seems likely.
>
> The underlying issue seems to be in the design of VNICs and etherstub.
> On a regular NIC, there's a native instance that has its own MAC
> address and thus can work normally when plumbed; no VNIC is required.
> On an etherstub, though, the native instance has no address, and VNICs
> appear to be required in order to provide MAC addresses needed for
> clients.
>
> Thus, to a client, the two are different.

That's correct.

> Providing etherstubs themselves with arbitrary MAC addresses and
> allowing them to be plumbed normally (just as though they were regular
> NICs) would resolve the issue, but might be a step too far.

That would be one option, I'm sure we can work something out to make  
etherstub bridging-friendly.

>>> I'm not sure, though, that it's an interesting case.  You'll get
>>> better performance if you just put all of the VNICs that must talk
>>> with each other together on a single etherstub if you're planning to
>>> bridge etherstubs together.  If you're planning to bridge an  
>>> etherstub
>>> with a regular NIC, then just move the VNICs over to the regular  
>>> NIC.
>>
>> An important benefit is to have the flexibility to build virtual
>> networks in a box which map directly to physical topologies.
>
> Please expand on that.

Our architecture allows one to build a virtual network in a box  
consisting of virtual switches, virtual NICs, etc. That virtual  
network in a box can be used to simulate physical network topologies,  
which can be used by testers, developers, deployers, etc. The more  
virtual network elements we have, the richer these virtual networks  
can be, and the more realistically they can be used to simulate real  
networks.

> If we're talking about using these for general testing purposes (e.g.,
> creating a four-port bridge inside a box and using that for the IPMP
> test suite; no external hardware required), then I think etherstubs
> are just the wrong mechanism, because they are so different from
> actual Ethernets.
>
> What I believe we really need to generalize that case is an entity
> that emulates a point-to-point link, where the two endpoints are
> separate MACs in the same box.  Such a pseudo-device would give us:
>
>  - Distinct transmit and receive paths across the emulated link,
>    which could be subject to any desired impairment (e.g., MTU
>    restrictions, loss rates, delay, link up/down).
>
>  - Emulation of negotiated L2 parameters, including all of the
>    standard Ethernet capabilities.
>
>  - An obvious means to test wireless configurations, by having one
>    endpoint behave as an 802.11 "client" and the other behave as an
>    AP, emulating the assocation/disassociation/authentication
>    mechanisms.
>
> None of that appears to be as feasible with etherstubs.

What you are describing should be possible using VNICs connected  
through etherstubs. But this is different than what I was referring  
to, which is a bridge between etherstubs.

> If we're talking about other cases, then I'll need more information.
> One case suggested by private email was a DomU migrating from one NIC
> to another, and using an etherstub bridged to the actual NIC in order
> to make the transition "easy."  But if that's the goal, I would
> suggest that having either a mechanism that allows migration of VNICs
> among NICs without teardown or something that inserts an indirection
> mechanism (perhaps even an etherstub with the optional capability of
> linking to an underlying NIC) would be better.

> Don't forget that NICs used in bridges _must_ be in promiscuous mode,
> and that this destroys a good bit of performance.  That alone makes
> them less interesting for ad-hoc system reconfiguration types of
> activities.

This is different, we've talked about the VNICs migration during the  
Crossbow design review. The design is there, but the feature was not  
implemented as part of the first Crossbow putback.

>>> With Crossbow, the classification is tied to the administrative  
>>> bits,
>>> which rely on explicit configuration of the VNICs and flows involved
>>> using a user-space component.  With bridging, forwarding entries are
>>> created and updated on the fly based on source MAC addresses seen in
>>> the data path, and then aged away over time; there's no  
>>> administrative
>>> involvement normally expected for these entries.
>>
>> This doesn't have to be the case. The Crossbow flow implementation
>> provides a kernel API which allows flows to be created and added to
>> flow tables. That API today is used for VNICs, but also via MAC  
>> client
>> creation in general (e.g. through LDOMs), for user-specified flows,
>> and for multicast addresses. It could be used by the bridge code as
>> well.
>
> The mac_flow_add() function asserts that the ft_mip perimeter is
> held.  My understanding of how the perimeters are used in Crossbow
> (please correct me if I've gotten this wrong) is that they're *NEVER*
> taken in the upward direction, as that would lead to deadlock.  This
> is given as general rule "R2" in the mac.c block comments.
>
> Assuming that to be correct, it would mean that potentially _all_
> received packets would have to be shuffled off to a separate kernel
> thread, so that the source MAC address could be safely inspected and
> used to update the list of learned MAC destinations.

That wouldn't be necessarily required for all packets. Flow additions  
and removals currently require the perimeter to be held, but not  
lookups of course.

> Worse still, it depends on a per-MAC lock for a per-MAC flow table,
> which makes no sense for bridging.  Bridge forwarding entries are a
> common database for a given bridge instance -- packets flowing in or
> out of any port on the bridge are matched against the common pool, not
> against a set of per-port entries.  We thus need a resource common to
> multiple MACs, which is not something that appears to be part of
> Crossbow, or we need to multiply the storage required by replicating
> each entry by N (and the locking effort goes up by N as well), so we
> can remain with Crossbow's per-port mi_flow_tab.

Yes, some of these entries will be duplicated in multiple mac_impl_t  
flow tables. On the other end the data-path can be kept simpler and  
only one classification is needed.

An alternative would be to have a per bridge flow table, like the per  
MAC client flow table we use for user-specified flows, and do a lookup  
in that table when forwarding is enabled, however this requires an  
additional lookup.

>>> I regard all of that as a research project.  It may well be an
>>> interesting one, but it's not this project by any stretch.  I have  
>>> no
>>> plans or engineering resources available to redesign the internals  
>>> of
>>> Crossbow to handle things it wasn't originally designed to do, and I
>>> think that insisting on such an extension of the project I've  
>>> proposed
>>> is not reasonable.  I will not be doing that.
>>
>> I don't think you need to "redesign the internals of Crossbow".
>>
>> We have kernel APIs which I believe can achieve most of what you need
>> here. There might be some small gaps, but I don't see why you would
>> need to introduce a new classification table at layer 2 since we
>> already have most of what you need at the same layer in mac.
>
> I've stated my needs previously, and I don't see how Crossbow's
> features would fill them or be modified to fill them.  To state them
> again in detail:
>
>  - I need to be able to add and delete entries in the table from
>    within the datapath; this means manipulating the flow table on an
>    upcall from the driver layer.  This means having a locking scheme
>    that supports table modification in an upward direction.

The flow lookup can be done from the data-path without holding the  
perimeter of course, but you are correct that the addition or removal  
of flow would have to be handed-off to a helper thread. Is there a  
particular reason why the flow update via helper thread would be  
problematic?

>  - I need to control the forwarding function on a per-port basis in
>    order to implement Spanning Tree.  When disabled (the default),
>    arriving packets are never matched against bridge forwarding
>    entries, and are just delivered to local destinations (i.e., VNICs
>    on that port).  When forwarding is enabled by STP, arriving
>    packets are matched against forwarding entries and delivered as
>    directed.  This means we need to match against subsets of flows
>    (local to port or all flows) depending on port state.

The separate flow table would probably be ideal here. An alternative  
would be to have the entries still in place, but not do the forwarding  
on a match if it is disabled.

>
>  - I need to do a special match on "unknown destinations" -- anything
>    not in the forwarding table must be copied to every port that's
>    enabled by the control described above.

An unknown destination can be easily handled, since it results in a  
NULL flow during a matching operation, and is not on the performance  
sensitive case.

>
>  - The forwarding table must be implemented in a per-bridge manner,
>    rather than per-port/link, as is currently done with Crossbow.
>    Forwarding is a global function for a bridge; otherwise, you can't
>    actually forward between ports.

As discussed above, we could replicate the flows in each table of the  
links that are part of a bridge, or we could have a separate flow  
table per bridge, separate from the per mac_impl_t flow tables.

>  - If I'm to support IVL and SVL, I need to vary my lookup so that it
>    sometimes matches on MAC address alone (SVL) and other times based
>    on MAC+VID (IVL).  (I could ditch the feature and hard code for
>    one or the other, but that'd be less good.)

We currently do matching on MAC + VID, but in previous implementations  
we had code which could toggle between MAC only or MAC + VID based  
classification. So we can certainly consider having a flexible  
approach which satisfies this requirement.

>  - I need to age away entries (potentially based on bridge
>    parameters).  (This part at least looks "easy" in that I would use
>    a separate kernel thread that wakes up when needed.)
>
>
> There are many other issues that I don't know how to resolve.  For a
> few of these:
>
>  - Crossbow flows are associated with resources (CPUs and the like),
>    and these are currently assigned when flows are created by
>    administrative action, but it's unclear what choices data-driven
>    flow creation should use for those parameters.  There's probably a
>    separate infrastructure needed here -- perhaps per-VLAN set of
>    flow 'templates' set up by an administrator that get used to
>    create the actual flows -- but it's unclear how that should work.

The resources are not used by the flows themselves directly but rather  
by the data scheduling entities which relies on the flows, e.g. the  
SRS. If you use the flows directly you can decide how you will use  
these resource parameters. You can decide to not do any resource  
management, or have per bridge resources, etc.

>  - What happens when L2 forwarding occurs?  Do we use the same input
>    side flow for the output side, or are two separate look-ups done?
>    Things are clearer for IP and other network layer features using
>    Crossbow, as the xb responsibility effectively ends at the network
>    layer, so the user naturally expects a new flow look-up if the
>    packet is transmitted elsewhere by IP.  It's unclear if that makes
>    sense for forwarding done entirely inside Crossbow.

I'm not sure what problem you are referring to. Are you referring to  
user flows managed through flowadm(1M)?

>
>   - Is "unknown destination" (transmit on all ports) a single flow or
>    N separate output-side look-ups?

If you don't find your destination we can then treat this as a special  
case which does the appropriate transmission(s). We already have such  
a special case when multiple MAC clients are present which causes the  
packet to go out on the wire.

>  - What happens when the user configures MAC based flows *and*
>    bridging is in use?  We then have two administrators manipulating
>    the same table -- the human admin is creating static entries, and
>    the system is creating and deleting dynamic ones.  What happens
>    when they overlap?

flowadm(1M) entries are L3 and up, so they wouldn't conflict with  
bridging. The only L2 entries we currently have are for MAC clients,  
and bridging already know how to handle these.

>>> I can see that this path wasn't taken, so I can't help but wonder  
>>> how
>>> reuse of Crossbow's classifier could be considered a requirement for
>>> bridging.
>>
>> It is very relevant to bridging since the bridge forwarding happens  
>> at
>> the same place on the data-path as the classification that Crossbow
>> introduced in the MAC layer. For example on transmit, the
>> classification on the destination MAC address results in sending the
>> packet to another MAC client (e.g. VNIC), send copies of the packets
>> to members of a multicast group, or send the packet on the wire. A  
>> new
>> outcome would be to pass the packet to a bridge.
>
> The flow identification in flow_ip_v4_match() does the same sort of
> look-up that's done by IPv4 forwarding entries, IP Filter, and IPsec
> policy matching.  That doesn't make Crossbow a replacement for any of
> those other cases, though.
>
> At a high enough level, the two functions do appear to be similar.
> They're not the same, though, and have significant differences when
> you look at the details.  For instance, the things I'm matching
> against are *not* per-port entries.

They are not today in your design, but that's not a requirement, the  
MAC addresses could be added to the per MAC instance flow table, or  
you could even have a separate flow table per bridge.

>> Similarly on receive, the old mac rx_add() entry points are gone, and
>> demultiplexing to the interested parties is now done by the mac layer
>> through the same classification table. So having the entry for an
>> address the bridge is interested in would allow the classification to
>> be leveraged for the receive side as well.
>>
>> So reusing the classifier for components of the same layer of the
>> stack seems that the natural thing to do. Once you use flows you can
>> also take advantage of hardware classification on the receive side.
>>
>> The Crossbow team will be happy to answer questions you may have  
>> about
>> the new datapath, and discuss specific requirements you may have, and
>> are not addressed by the current implementation.
>
> Again, I have no plan or resources available to launch the necessary
> research project into the redesign of Crossbow for the use of
> bridging.  It sounds like an interesting thing, to be sure, but that's
> just not going to happen within the context of this project.

We're not talking about a research project, and we're not talking  
about redesigning Crossbow. Crossbow introduced a new framework in  
place in Solaris to do flow classification which enables steering of  
packets based on layer-2 addresses, which corresponds to one of the  
needs of bridging.

> If you feel a TCR or outright denial of this project is necessary to
> preserve Crossbow, then speak with the other ARC members and build a
> consensus for that.

I don't think we're are talking about "preserving Crossbow". This  
discussion is about using a common framework to implement a new  
feature. In the long term it will make the framework itself more  
complete, allow bridging to take advantage of future improvements,  
hardware classification, and avoid duplication.

PSARC members should now have the information needed to make an  
informed decision. I'll be happy to provide more information if needed.

> If someone else has the necessary resources to do the required
> research (or even wants to launch a counter-project), then be my
> guest.
>
>>> One important issue did come up here: we need to define the relative
>>> ordering between L2 filtering and bridging, and I believe it makes
>>> sense to put L2 filtering closer to the physical I/O.  In other  
>>> words,
>>> L2 filter should do its work underneath the bridge.
>>
>> There's filtering which needs to occur between multiple MAC clients
>> (VNICs are MAC clients) defined on top of the same data-link. For
>> example to be consistent with the way things work in the physical
>> world, one might want to prevent a VM to be able to specific send
>> packets on the wire, which in this case would include a bridge. On  
>> the
>> transmit side these checks would have to be done before the packet is
>> potentially sent through a bridge, i.e. the L2 filtering would have  
>> to
>> be done "on top" of the bridge.
>
> The issue that I'm pointing out is largely an administrative one.
> When someone creates an L2 filter, what exactly are they expecting to
> filter on?
>
> If they're expecting to filter on the actual physical interface that
> bridging uses, then the scheme you're suggesting won't work right --
> the action of bridge forwarding can (and does) redirect output packets
> from the originally intended link over to the actual one where the
> destination resides -- or to all ports if the destination is not
> known.  If an administrator says "never send packets to
> 00:01:02:03:04:05 on hme0", then he'll be quite surprised to find that
> when he sends packets to that address on ce0, the packet comes
> stumbling out hme0 despite his filter, because the L2 filtering "on
> top" of the bridge saw only ce0 as the output.

The primary goal is the filtering on a per MAC client basis. I.e. when  
the users specifies "hme0", this corresponds to the traffic going  
through the primary MAC client of hme0, for example IP on top of that  
data-link. It is different from all traffic going through the physical  
MAC instance which is shared by multiple MAC clients/VNICs/etc.

> Perhaps a hybrid approach is needed.  Place the hooks in two places:
> below the bridge for physical I/O and inside the MAC layer for
> client-to-client (VNIC-to-VNIC) communication.  The former will allow
> the user to filter packets with reference to the actual physical
> hardware in use (as opposed to just what the network layer "thinks" is
> in use), and allow him to filter traffic between MAC clients as well.
> The latter would not involve bridging, and thus would use the link as
> known to the network layer.

L2 Filtering based on "physical" interfaces is not currently planned,  
but could be added in the future if needed.

Nicolas.

>
>
> -- 
> James Carlson, Solaris Networking              <james.d.carlson@sun.com 
> >
> Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442  
> 2084
> MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442  
> 1677

-- 
Nicolas Droux - Solaris Kernel Networking - Sun Microsystems, Inc.
nicolas.droux@sun.com - http://blogs.sun.com/droux


From carlsonj@phorcys.east.sun.com Wed Dec 24 07:48:03 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBOFm3Qr021997
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 24 Dec 2008 07:48:03 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mBOFm2KY062335;
	Wed, 24 Dec 2008 08:48:02 -0700 (MST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KCE00907181OO00@brm-avmta-1.central.sun.com>; Wed,
 24 Dec 2008 08:48:01 -0700 (MST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KCE003L11801140@brm-avmta-1.central.sun.com>; Wed,
 24 Dec 2008 08:48:01 -0700 (MST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id mBOFm0De021956; Wed,
 24 Dec 2008 10:48:00 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id mBOFm0pS021953; Wed,
 24 Dec 2008 10:48:00 -0500 (EST)
Date: Wed, 24 Dec 2008 10:48:00 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: Issues for 2008/055
In-reply-to: <1491995A-D70F-4856-9627-36EC11B0B8CF@sun.com>
To: Nicolas Droux <droux@sun.com>
Cc: Nicolas Droux <Nicolas.Droux@sun.com>, PSARC-ext@sun.com
Message-id: <18770.22960.710148.605664@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
 <18761.26709.246151.818319@gargle.gargle.HOWL>
 <FEF6920A-4380-46B8-96C1-2D9E471C5C12@Sun.COM>
 <18762.31120.53165.970797@gargle.gargle.HOWL>
 <1491995A-D70F-4856-9627-36EC11B0B8CF@sun.com>
Status: RO
Content-Length: 19732

Nicolas Droux writes:
> On Dec 18, 2008, at 9:25 AM, James Carlson wrote:
> > Providing etherstubs themselves with arbitrary MAC addresses and
> > allowing them to be plumbed normally (just as though they were regular
> > NICs) would resolve the issue, but might be a step too far.
> 
> That would be one option, I'm sure we can work something out to make  
> etherstub bridging-friendly.

OK.  Let me know when that happens.

For now, I'm planning to detect the case of etherstubs and handling
them as a special case: no STP (and thus no DLPI needed) and
forwarding enabled when configured.

> >> An important benefit is to have the flexibility to build virtual
> >> networks in a box which map directly to physical topologies.
> >
> > Please expand on that.
> 
> Our architecture allows one to build a virtual network in a box  
> consisting of virtual switches, virtual NICs, etc. That virtual  
> network in a box can be used to simulate physical network topologies,  
> which can be used by testers, developers, deployers, etc. The more  
> virtual network elements we have, the richer these virtual networks  
> can be, and the more realistically they can be used to simulate real  
> networks.

Except that the virtual environments constructed this way do not
emulate the real ones faithfully.

For instance (and it's just one instance), there's no 802 parameter
negotiation, so if the testing desired involves anything related to
that -- such as duplex or speed, both of which affect aggregations and
bridging -- then VNICs and etherstubs are right out.

In short, no, that doesn't solve the problem well.  It does work for
some cases -- tests that involve only higher level protocols, and then
likely a subset of those that don't need to see hardware-like behavior
-- but doesn't for many others.  That's why I proposed point-to-point
emulations in addition to etherstubs.  They'd provide a better way to
construct intentional test environments.

> > None of that appears to be as feasible with etherstubs.
> 
> What you are describing should be possible using VNICs connected  
> through etherstubs.

I'd like to see it done.  In the meantime, I'll use other means to
test.

> But this is different than what I was referring  
> to, which is a bridge between etherstubs.

Yes ... but I was trying to fish out the reason for such a bridge.  It
seems hard to understand.

> > Don't forget that NICs used in bridges _must_ be in promiscuous mode,
> > and that this destroys a good bit of performance.  That alone makes
> > them less interesting for ad-hoc system reconfiguration types of
> > activities.
> 
> This is different, we've talked about the VNICs migration during the  
> Crossbow design review. The design is there, but the feature was not  
> implemented as part of the first Crossbow putback.

OK.  Someone mentioned the possible usage of bridges for migration in
a private discussion, and I brought it up here to find out what sorts
of concrete situations you might have been referring to in this way:

   An important benefit is to have the flexibility to build virtual
   networks in a box which map directly to physical topologies.

I don't know what that means, so I'm searching for concrete answers.
If it means trying to use VNICs on etherstubs as a way to run tests,
then, as I've pointed out, that only works for some situations, and
doesn't for others.  Rather pointedly, I can't test bridging that way,
because I need Ethernet NICs.

> > Assuming that to be correct, it would mean that potentially _all_
> > received packets would have to be shuffled off to a separate kernel
> > thread, so that the source MAC address could be safely inspected and
> > used to update the list of learned MAC destinations.
> 
> That wouldn't be necessarily required for all packets. Flow additions  
> and removals currently require the perimeter to be held, but not  
> lookups of course.

It gets worse.

We have to do a look-up first on source address, and then on
destination.  That source look-up isn't something that Crossbow does
today, so that's a new addition.

If the source look-up *either* finds no entry *or* finds an entry that
points to a different output, then we have to make changes.  We must
delete the old entry (if present) and create a new one.

We will be forced (by Crossbow architecture) to put that on a separate
thread for processing, which means that new packets can arrive while
we're trying to do that work.  Those new packets must also be queued
until the flow is ready.

Once we get the new flow inserted, our troubles do not end.  The 802
standards require that we provide in-order delivery to
"conversations."  Since we don't have per-MAC-pair storage, we don't
know what "conversations" exist, which means that we need to preserve
order among things that we're forced to queue, at least for a given
input port.

In other words, as long as there are still packets queued in this new
mechanism for a given source address, ones that have not yet been
drained following the flow insertion, we have to enqueue any
subsequent packets behind them.  Only when the queue empties can we
return to normal Crossbow processing and avoid reordering.

How do we do this?  It can't be a reference count on the flow, because
it was the non-existence of a usable flow that would have caused us to
go through this task-based detour in the first place.  It must be some
sort of new structure, but I'm not sure what it looks like.  It'll
take at least a bit of research to figure out how to put this
together.

> > Worse still, it depends on a per-MAC lock for a per-MAC flow table,
> > which makes no sense for bridging.  Bridge forwarding entries are a
> > common database for a given bridge instance -- packets flowing in or
> > out of any port on the bridge are matched against the common pool, not
> > against a set of per-port entries.  We thus need a resource common to
> > multiple MACs, which is not something that appears to be part of
> > Crossbow, or we need to multiply the storage required by replicating
> > each entry by N (and the locking effort goes up by N as well), so we
> > can remain with Crossbow's per-port mi_flow_tab.
> 
> Yes, some of these entries will be duplicated in multiple mac_impl_t  
> flow tables.

No, not "some."  "All."  It's a requirement of bridging that
forwarding is the same no matter which input port is used.

The implication is NxM storage for forwarding entries (N links, M
known MAC destinations), each represented as a separate flow, and N
times as much effort in inserting or deleting entries, which can
easily happen en masse when topology changes occur.

A normal occurance is that a port goes "up" somewhere in the network,
and you start learning the full set of MAC addresses through a
different port.  Each one of those will require a separate flow delete
and insert operation on each one of the ports, and queuing of all the
traffic as this happens.  (With behavior much like the old outer
STREAMS perimeter in IP, I suspect.)

But -- wait -- it gets worse.

Those flows are distinct units.  We only match one.  But the real
usefulness in Crossbow is in being able to manage flows
administratively, and I still don't see how that can work when
bridging creates and destroys flows on its own.  Should the user just
be prohibited for all time from controlling flows through the bridge?

> On the other end the data-path can be kept simpler and  
> only one classification is needed.
> 
> An alternative would be to have a per bridge flow table, like the per  
> MAC client flow table we use for user-specified flows, and do a lookup  
> in that table when forwarding is enabled, however this requires an  
> additional lookup.

Yes.  And it also requires some definition of what occurs when flows
are matched in more than one place.

That's actually a deeper problem anyway with this whole scheme, as
you'd be matching a flow on input and then (assuming administrative
controls) another one on output, and it's not clear how they should
work together.  What if they specify conflicting resource
restrictions?

With IP, the usage is much more obvious.  When a packet matches a flow
and is delivered to IP, the "flowness" of the packet stops there.  If
IP decides to forward the packet out a different interface, the MAC
layer will do a brand new ex nihilo look-up for a new flow to match in
that context.  The original horse it rode in on doesn't count.

Does the same apply for bridging?  It's really unclear.

> >  - I need to be able to add and delete entries in the table from
> >    within the datapath; this means manipulating the flow table on an
> >    upcall from the driver layer.  This means having a locking scheme
> >    that supports table modification in an upward direction.
> 
> The flow lookup can be done from the data-path without holding the  
> perimeter of course, but you are correct that the addition or removal  
> of flow would have to be handed-off to a helper thread. Is there a  
> particular reason why the flow update via helper thread would be  
> problematic?

Yes.  Because that forces data packets in a very common case (ordinary
bridge learning) into a slow and highly complex queuing mechanism.
The design and properties of such a mechanism are well outside the
scope of this project, and are *purely* needed to cope with Crossbow
design peculiarities -- it has nothing to do with bridging, and
everything to do with getting around rule "R2."

> >  - I need to control the forwarding function on a per-port basis in
> >    order to implement Spanning Tree.  When disabled (the default),
> >    arriving packets are never matched against bridge forwarding
> >    entries, and are just delivered to local destinations (i.e., VNICs
> >    on that port).  When forwarding is enabled by STP, arriving
> >    packets are matched against forwarding entries and delivered as
> >    directed.  This means we need to match against subsets of flows
> >    (local to port or all flows) depending on port state.
> 
> The separate flow table would probably be ideal here.

Yes.  However, I think it'd likely involve at least a minor
restructuring of Crossbow to get there.  Something would have to know
when to invoke searches on that new table.

> An alternative  
> would be to have the entries still in place, but not do the forwarding  
> on a match if it is disabled.

That's part of it.  The rest is that the port still needs to function
as a normal port when forwarding isn't active, which means that
receive to local destinations still works and transmit out that one
port (at least for STP) still works.

> >  - I need to do a special match on "unknown destinations" -- anything
> >    not in the forwarding table must be copied to every port that's
> >    enabled by the control described above.
> 
> An unknown destination can be easily handled, since it results in a  
> NULL flow during a matching operation, and is not on the performance  
> sensitive case.

It involves modifying Crossbow to do something special with that NULL
case.

> >  - Crossbow flows are associated with resources (CPUs and the like),
> >    and these are currently assigned when flows are created by
> >    administrative action, but it's unclear what choices data-driven
> >    flow creation should use for those parameters.  There's probably a
> >    separate infrastructure needed here -- perhaps per-VLAN set of
> >    flow 'templates' set up by an administrator that get used to
> >    create the actual flows -- but it's unclear how that should work.
> 
> The resources are not used by the flows themselves directly but rather  
> by the data scheduling entities which relies on the flows, e.g. the  
> SRS. If you use the flows directly you can decide how you will use  
> these resource parameters. You can decide to not do any resource  
> management, or have per bridge resources, etc.

Right.  And as I'm saying, I don't quite know what to do with this.

If I do nothing, and bridge flows have no resource controls, then I
don't see that I've given the user anything noteworthy by using
Crossbow than I do by implementing my own forwarding.  There are no
extra features that become usable, but there's a whole lot more
complexity in the implementation, and much more interesting failure
modes that can result.

The forwarding look-up process itself is trivial.  I'm using the
existing kernel AVL trees to do the work for me.  Trying to reuse
Crossbow's flow matching mechanism as though it were a bridge
forwarding process involves adding substantial complexity, and I'm
just not seeing any obvious benefit.

> >  - What happens when L2 forwarding occurs?  Do we use the same input
> >    side flow for the output side, or are two separate look-ups done?
> >    Things are clearer for IP and other network layer features using
> >    Crossbow, as the xb responsibility effectively ends at the network
> >    layer, so the user naturally expects a new flow look-up if the
> >    packet is transmitted elsewhere by IP.  It's unclear if that makes
> >    sense for forwarding done entirely inside Crossbow.
> 
> I'm not sure what problem you are referring to. Are you referring to  
> user flows managed through flowadm(1M)?

Yes, in part.

For IP, the situation is clear: input and output are independent.  For
bridging, it's less clear.  Output is an artifact or outcome of having
done the input-side flow identification.

Does it make sense to do output-side flow look-ups with bridging?
They wouldn't be needed to do anything that a bridge needs to do, but
the administrative model of Crossbow would become very odd if output
flow controls worked only "sometimes."

> >   - Is "unknown destination" (transmit on all ports) a single flow or
> >    N separate output-side look-ups?
> 
> If you don't find your destination we can then treat this as a special  
> case which does the appropriate transmission(s). We already have such  
> a special case when multiple MAC clients are present which causes the  
> packet to go out on the wire.

That question was actually about the output-side flows.

> >  - What happens when the user configures MAC based flows *and*
> >    bridging is in use?  We then have two administrators manipulating
> >    the same table -- the human admin is creating static entries, and
> >    the system is creating and deleting dynamic ones.  What happens
> >    when they overlap?
> 
> flowadm(1M) entries are L3 and up, so they wouldn't conflict with  
> bridging. The only L2 entries we currently have are for MAC clients,  
> and bridging already know how to handle these.

A main point in using Crossbow would be to enable the administrative
mechanisms.

If we leave that behind, what is there?  A MAC-based search function?

> > At a high enough level, the two functions do appear to be similar.
> > They're not the same, though, and have significant differences when
> > you look at the details.  For instance, the things I'm matching
> > against are *not* per-port entries.
> 
> They are not today in your design, but that's not a requirement, the  
> MAC addresses could be added to the per MAC instance flow table, or  
> you could even have a separate flow table per bridge.

Quite simply put, this project is not rewriting Crossbow to provide
features that would allow the sort of design you've described.

If someone else wants to do that, then that's great.  If someone wants
to provide Crossbow-based interfaces that work well for bridging, then
let us know, and we'll see how to schedule a follow-on project to use
those.

That's not this project, and I refuse to allow this project to be made
dependent on Crossbow features that do not exist and that nobody is
working on.

> > Again, I have no plan or resources available to launch the necessary
> > research project into the redesign of Crossbow for the use of
> > bridging.  It sounds like an interesting thing, to be sure, but that's
> > just not going to happen within the context of this project.
> 
> We're not talking about a research project, and we're not talking  
> about redesigning Crossbow. Crossbow introduced a new framework in  
> place in Solaris to do flow classification which enables steering of  
> packets based on layer-2 addresses, which corresponds to one of the  
> needs of bridging.

As described many times over now, it does not fit the needs of
bridging.  In my opinion -- which I think counts as delivery of this
project is my responsibility -- it can't be made to fit the needs of
bridging without substantial and risky rework of internal elements of
Crossbow.

Just because you've got a hammer, that doesn't make every problem look
like a nail.  I'm not using that flow classification because it
doesn't fit.

> > If you feel a TCR or outright denial of this project is necessary to
> > preserve Crossbow, then speak with the other ARC members and build a
> > consensus for that.
> 
> I don't think we're are talking about "preserving Crossbow". This  
> discussion is about using a common framework to implement a new  
> feature. In the long term it will make the framework itself more  
> complete, allow bridging to take advantage of future improvements,  
> hardware classification, and avoid duplication.

It doesn't appear to do much of any of that.  We've already ruled out
flow administration at this level as an unclear concept, the hardware
features can't be used due to the use of both promiscuous mode and
searching based on source MAC address (the hard part of all this,
which the hardware doesn't do), and the "duplication" (if any) is in
the most trivial of places -- the table look-up function on
destination address, for which I use existing kernel facilities, and
thus don't actually duplicate anything.

> PSARC members should now have the information needed to make an  
> informed decision. I'll be happy to provide more information if needed.

I agree.

PSARC members: please vote to deny if you believe that bridging must
be built using Crossbow's classifier.  That's not this project, and
it's not going to be this project.  It's someone else's project.

> > If they're expecting to filter on the actual physical interface that
> > bridging uses, then the scheme you're suggesting won't work right --
> > the action of bridge forwarding can (and does) redirect output packets
> > from the originally intended link over to the actual one where the
> > destination resides -- or to all ports if the destination is not
> > known.  If an administrator says "never send packets to
> > 00:01:02:03:04:05 on hme0", then he'll be quite surprised to find that
> > when he sends packets to that address on ce0, the packet comes
> > stumbling out hme0 despite his filter, because the L2 filtering "on
> > top" of the bridge saw only ce0 as the output.
> 
> The primary goal is the filtering on a per MAC client basis. I.e. when  
> the users specifies "hme0", this corresponds to the traffic going  
> through the primary MAC client of hme0, for example IP on top of that  
> data-link. It is different from all traffic going through the physical  
> MAC instance which is shared by multiple MAC clients/VNICs/etc.

OK.  As long as the designers of L2 filtering can describe things so
that the user understands that it's on an internal MAC client basis,
rather than related to the physical port, that sounds fine to me.  I
would think it's confusing, but there are clear usage scenarios in
either direction, so I don't actually care which one they implement.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From droux@sun.com Sat Jan 31 16:17:08 2009
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n110H8qe007260
	for <psarc-ext@sac.sfbay.sun.com>; Sat, 31 Jan 2009 16:17:08 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id n110H6LJ008768
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Sat, 31 Jan 2009 17:17:08 -0700 (MST)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KED0040J24JO900@nwk-avmta-1.sfbay.Sun.COM> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Sat, 31 Jan 2009 16:17:07 -0800 (PST)
Received: from brmea-mail-1.sun.com ([192.18.98.31])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KED00B9124ISA40@nwk-avmta-1.sfbay.Sun.COM> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Sat,
 31 Jan 2009 16:17:06 -0800 (PST)
Received: from fe-amer-09.sun.com ([192.18.109.79])
	by brmea-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id n110H6ea020352	for
 <PSARC-ext@sun.com>; Sun, 01 Feb 2009 00:17:06 +0000 (GMT)
Received: from conversion-daemon.mail-amer.sun.com by mail-amer.sun.com
 (Sun Java(tm) System Messaging Server 7.0-3.01 64bit (built Dec  9 2008))
 id <0KED0070021J8X00@mail-amer.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Sat, 31 Jan 2009 17:17:06 -0700 (MST)
Received: from [10.0.0.2] ([unknown] [71.210.196.244])
 by mail-amer.sun.com (Sun Java(tm) System Messaging Server 7.0-3.01 64bit
 (built Dec  9 2008)) with ESMTPSA id <0KED0070K24HRL00@mail-amer.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Sat,
 31 Jan 2009 17:17:06 -0700 (MST)
Date: Sat, 31 Jan 2009 17:17:04 -0700
From: Nicolas Droux <droux@sun.com>
Subject: Re: Issues for 2008/055
In-reply-to: <18770.22960.710148.605664@gargle.gargle.HOWL>
Sender: Nicolas.Droux@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: PSARC-ext@sun.com
Message-id: <F9739781-85EB-491A-8D6A-B07ABD7F09DC@sun.com>
MIME-version: 1.0
X-Mailer: Apple Mail (2.930.3)
Content-type: text/plain; format=flowed; charset=US-ASCII
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
 <18761.26709.246151.818319@gargle.gargle.HOWL>
 <FEF6920A-4380-46B8-96C1-2D9E471C5C12@Sun.COM>
 <18762.31120.53165.970797@gargle.gargle.HOWL>
 <1491995A-D70F-4856-9627-36EC11B0B8CF@sun.com>
 <18770.22960.710148.605664@gargle.gargle.HOWL>
Status: RO
Content-Length: 2835

Jim,

I have included below a proposal which would allow bridging to
cleanly integrate with mac without the introduction of special
cases in the common data-path, and enabling bridging to
leverage the layer-2 classification that was introduced
by Crossbow. The proposal takes into account the issues we
discussed earlier in this thread.

- per bridge table

	- based on existing bridging implementation, captures MAC
	  addresses associated with ports, time stamps needed
	  to track lifetime of entries, etc.
	- table could be a flow table, but this is not the main
	  goal of this proposal.

- new promiscuous callback flag, synchronous
	
	- bridge registers promiscuous callback with mac,
	  specifies (new) sync flag
	- mac calls the registered callback with original packet
	  (read-only)
	- bridge inspects packet to extract MAC addresses
	- bridge updates its table according to MAC address
	- bridge calls mac to add entries to the layer-2 flow tables
	  of the ports associated with the bridge.

- updates to the MAC layer-2 flows

	- entries are added and removed to and from the layer-2
	  classification table of a MAC instance (port) via a flow
	  API provided by mac to the bridging code.
	- updates to these MAC classification tables are done
	  when addresses are added to/removed from the bridge table
	- The callback function is a bridge processing function,
	  the cookie points to the bridge table entry for the
	  MAC address
	- The flow addition/deletion API can be restructured
	  to no longer require the MAC perimeter to be held
	  when flow entries are added and deleted by the bridge.
	  This would allow the updates to be done without the
	  help of a worker thread, and would scale better with
	  frequent updates.
	- The flow API can be extended to allow a set of flows
	  to be removed with a single call. A MAC client handle
	  and cookie could be used to identify the set of
	  flows to be removed.

- TX/RX data path

	- does not require special processing on data-path
	- RX and TX classification causes the MAC addresses
	  registered by the bridge to be matched against
	  the destination MAC address of sent and received packets
	- packets are passed to bridge callback along with
	  registered cookie
	- bridge gets packet, and knows destination to forward packet
	  to appropriate destination associated with the MAC address
	- in the future, can takes advantage of TX fastpath
	  transparently
	- introduce a new "no match" mac callback, invoked when there's
	  no match on a MAC address lookup. Currently this is hard-coded
	  in the data-path to send the packet on the wire, and
	  a more generic mechanism would allow a bridge to receive
	  a copy of such packets.

Nicolas.

-- 
Nicolas Droux - Solaris Kernel Networking - Sun Microsystems, Inc.
droux@sun.com - http://blogs.sun.com/droux


From carlsonj@phorcys.east.sun.com Fri Feb 13 11:35:06 2009
Received: from sunmail2sca.sfbay.sun.com (sunmail2sca.SFBay.Sun.COM [129.145.155.234])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n1DJZ5Xd000189
	for <psarc-ext@sac.sfbay.sun.com>; Fri, 13 Feb 2009 11:35:06 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail2sca.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id n1DJYwGf028030;
	Fri, 13 Feb 2009 11:35:05 -0800 (PST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KF00081BRQGQL00@brm-avmta-1.central.sun.com>; Fri,
 13 Feb 2009 12:35:04 -0700 (MST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KF000EUQRQDR8B0@brm-avmta-1.central.sun.com>; Fri,
 13 Feb 2009 12:35:01 -0700 (MST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id n1DJYwNh012146; Fri,
 13 Feb 2009 14:34:58 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id n1DJYwQC012143; Fri,
 13 Feb 2009 14:34:58 -0500 (EST)
Date: Fri, 13 Feb 2009 14:34:58 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: Issues for 2008/055
In-reply-to: <F9739781-85EB-491A-8D6A-B07ABD7F09DC@sun.com>
To: Nicolas Droux <droux@sun.com>
Cc: PSARC-ext@sun.com
Message-id: <18837.52066.852097.442555@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
 <18761.26709.246151.818319@gargle.gargle.HOWL>
 <FEF6920A-4380-46B8-96C1-2D9E471C5C12@Sun.COM>
 <18762.31120.53165.970797@gargle.gargle.HOWL>
 <1491995A-D70F-4856-9627-36EC11B0B8CF@sun.com>
 <18770.22960.710148.605664@gargle.gargle.HOWL>
 <F9739781-85EB-491A-8D6A-B07ABD7F09DC@sun.com>
Status: RO
Content-Length: 1642

Nicolas Droux writes:
> I have included below a proposal which would allow bridging to
> cleanly integrate with mac without the introduction of special
> cases in the common data-path, and enabling bridging to
> leverage the layer-2 classification that was introduced
> by Crossbow. The proposal takes into account the issues we
> discussed earlier in this thread.

Thanks.  I think the high-level issue here is deciding on a way
forward, much of which is probably outside of the architectural review
realm.  The issues I see from the bridging project side of things are:

  - Is the project to redesign the Crossbow locking structures and
    otherwise provide bridging-friendly features something that's
    funded?  If so, what's its schedule?  If not, then are you
    actually suggesting that the bridging project be made dependent on
    something that isn't planned?  It would be at least a bit silly
    for me to agree to something like that.

  - How do we work towards a common understanding of the goals and
    requirements?  I'm not sure this list is the right place to do
    that, nor do I know who should "own" this task.  (Does it belong
    to some future project?)

I'll follow up with a detailed technical response to this proposal on
networking-discuss, as I _think_ we're no longer talking about the
bridging project that's under ARC review, but rather talking about a
potential future project.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Sebastien.Roy@sun.com Fri Feb 13 13:16:13 2009
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n1DLGCZP025425
	for <psarc-ext@sac.sfbay.sun.com>; Fri, 13 Feb 2009 13:16:12 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id n1DLFvIt004202
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Fri, 13 Feb 2009 21:16:11 GMT
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KF000803WEWZL00@nwk-avmta-2.sfbay.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Fri, 13 Feb 2009 13:16:08 -0800 (PST)
Received: from brmea-mail-2.sun.com ([192.18.98.43])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KF00080KWEV9P10@nwk-avmta-2.sfbay.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Fri,
 13 Feb 2009 13:16:08 -0800 (PST)
Received: from fe-amer-10.sun.com ([192.18.109.80])
	by brmea-mail-2.sun.com (8.13.6+Sun/8.12.9) with ESMTP id n1DLG7dt004584	for
 <PSARC-ext@sun.com>; Fri, 13 Feb 2009 21:16:07 +0000 (GMT)
Received: from conversion-daemon.mail-amer.sun.com by mail-amer.sun.com
 (Sun Java(tm) System Messaging Server 7.0-3.01 64bit (built Dec 23 2008))
 id <0KF000B00W7YHK00@mail-amer.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Fri, 13 Feb 2009 14:16:07 -0700 (MST)
Received: from [129.148.174.103] ([unknown] [129.148.174.103])
 by mail-amer.sun.com
 (Sun Java(tm) System Messaging Server 7.0-3.01 64bit (built Dec 23 2008))
 with ESMTPSA id <0KF0000YMWEVJHB0@mail-amer.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Fri, 13 Feb 2009 14:16:07 -0700 (MST)
Date: Fri, 13 Feb 2009 16:15:53 -0500
From: Sebastien Roy <Sebastien.Roy@sun.com>
Subject: Re: Issues for 2008/055
In-reply-to: <18837.52066.852097.442555@gargle.gargle.HOWL>
Sender: Sebastien.Roy@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: Nicolas Droux <droux@sun.com>, PSARC-ext@sun.com
Message-id: <1234559753.22408.108.camel@strat>
Organization: Sun Microsystems
MIME-version: 1.0
X-Mailer: Evolution 2.24.2
Content-type: text/plain; charset=UTF-8
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
 <18761.26709.246151.818319@gargle.gargle.HOWL>
 <FEF6920A-4380-46B8-96C1-2D9E471C5C12@Sun.COM>
 <18762.31120.53165.970797@gargle.gargle.HOWL>
 <1491995A-D70F-4856-9627-36EC11B0B8CF@sun.com>
 <18770.22960.710148.605664@gargle.gargle.HOWL>
 <F9739781-85EB-491A-8D6A-B07ABD7F09DC@sun.com>
 <18837.52066.852097.442555@gargle.gargle.HOWL>
Status: RO
Content-Length: 1535

On Fri, 2009-02-13 at 14:34 -0500, James Carlson wrote:
> Nicolas Droux writes:
> > I have included below a proposal which would allow bridging to
> > cleanly integrate with mac without the introduction of special
> > cases in the common data-path, and enabling bridging to
> > leverage the layer-2 classification that was introduced
> > by Crossbow. The proposal takes into account the issues we
> > discussed earlier in this thread.
> 
> Thanks.  I think the high-level issue here is deciding on a way
> forward, much of which is probably outside of the architectural review
> realm.

What is architecturally relevant, I think, is the feature intersection
of administratively-defined flows (and the things that come with them
like resource management and accounting) and bridging.  For example,
giving the administrator the ability to do flow accounting on a bridged
link is conceivable.  From a high-level administrative view, it's
something that I might expect would just work given the tools given to
me (dladm and flowadm) unless documented otherwise.

To me, this case is complete as-is as long as we understand and document
where such features don't interact.  If the architecture of this case
were incompatible with future work aimed at improving such feature
interaction, then I'd feel differently, but I don't see that this is
true.

If we can agree that such work is indeed in the realm of a future
project, then I'd suggest including advice to that effect to the
project's management in the opinion for this case.

-Seb



From droux@sun.com Thu Feb 19 11:36:53 2009
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n1JJaqJC015642
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 19 Feb 2009 11:36:53 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id n1JJan2E018713
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Thu, 19 Feb 2009 19:36:51 GMT
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KFB0000TVTEPO00@nwk-avmta-1.sfbay.Sun.COM> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Thu, 19 Feb 2009 11:36:50 -0800 (PST)
Received: from brmea-mail-1.sun.com ([192.18.98.31])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KFB00BQ4VTC0TC0@nwk-avmta-1.sfbay.Sun.COM> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Thu,
 19 Feb 2009 11:36:48 -0800 (PST)
Received: from fe-amer-09.sun.com ([192.18.109.79])
	by brmea-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id n1JJamUu028279	for
 <PSARC-ext@sun.com>; Thu, 19 Feb 2009 19:36:48 +0000 (GMT)
Received: from conversion-daemon.mail-amer.sun.com by mail-amer.sun.com
 (Sun Java(tm) System Messaging Server 7.0-3.01 64bit (built Dec 23 2008))
 id <0KFB00M00VNDAM00@mail-amer.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Thu, 19 Feb 2009 12:36:48 -0700 (MST)
Received: from [10.0.0.2] ([unknown] [129.150.18.180])
 by mail-amer.sun.com (Sun Java(tm) System Messaging Server 7.0-3.01 64bit
 (built Dec 23 2008)) with ESMTPSA id <0KFB004O6VT9EZ00@mail-amer.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Thu,
 19 Feb 2009 12:36:47 -0700 (MST)
Date: Thu, 19 Feb 2009 12:36:44 -0700
From: Nicolas Droux <droux@sun.com>
Subject: Re: Issues for 2008/055
In-reply-to: <1234559753.22408.108.camel@strat>
Sender: Nicolas.Droux@sun.com
To: Sebastien Roy <Sebastien.Roy@sun.com>
Cc: James Carlson <James.D.Carlson@sun.com>, PSARC-ext@sun.com
Message-id: <BF0EFF92-F807-4FD5-BD69-6FAF60A3DDC5@sun.com>
MIME-version: 1.0
X-Mailer: Apple Mail (2.930.3)
Content-type: text/plain; delsp=yes; format=flowed; charset=US-ASCII
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
 <18761.26709.246151.818319@gargle.gargle.HOWL>
 <FEF6920A-4380-46B8-96C1-2D9E471C5C12@Sun.COM>
 <18762.31120.53165.970797@gargle.gargle.HOWL>
 <1491995A-D70F-4856-9627-36EC11B0B8CF@sun.com>
 <18770.22960.710148.605664@gargle.gargle.HOWL>
 <F9739781-85EB-491A-8D6A-B07ABD7F09DC@sun.com>
 <18837.52066.852097.442555@gargle.gargle.HOWL>
 <1234559753.22408.108.camel@strat>
Status: RO
Content-Length: 2093


On Feb 13, 2009, at 2:15 PM, Sebastien Roy wrote:

> On Fri, 2009-02-13 at 14:34 -0500, James Carlson wrote:
>> Nicolas Droux writes:
>>> I have included below a proposal which would allow bridging to
>>> cleanly integrate with mac without the introduction of special
>>> cases in the common data-path, and enabling bridging to
>>> leverage the layer-2 classification that was introduced
>>> by Crossbow. The proposal takes into account the issues we
>>> discussed earlier in this thread.
>>
>> Thanks.  I think the high-level issue here is deciding on a way
>> forward, much of which is probably outside of the architectural  
>> review
>> realm.
>
> What is architecturally relevant, I think, is the feature intersection
> of administratively-defined flows (and the things that come with them
> like resource management and accounting) and bridging.  For example,
> giving the administrator the ability to do flow accounting on a  
> bridged
> link is conceivable.  From a high-level administrative view, it's
> something that I might expect would just work given the tools given to
> me (dladm and flowadm) unless documented otherwise.

Yes, I don't think there are any issues here.

> To me, this case is complete as-is as long as we understand and  
> document
> where such features don't interact.  If the architecture of this case
> were incompatible with future work aimed at improving such feature
> interaction, then I'd feel differently, but I don't see that this is
> true.
>
> If we can agree that such work is indeed in the realm of a future
> project, then I'd suggest including advice to that effect to the
> project's management in the opinion for this case.

I agree with that approach. For the short term we will look into  
providing a mechanism that will allow subsystems like bridging and  
layer-2 filtering to intercept packets without requiring the addition  
of checks specific to these subsystems in the common mac TX and RX  
data-paths.

Nicolas.

-- 
Nicolas Droux - Solaris Kernel Networking - Sun Microsystems, Inc.
droux@sun.com - http://blogs.sun.com/droux


From carlsonj@phorcys.east.sun.com Thu Feb 19 11:56:44 2009
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n1JJuh1K017143
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 19 Feb 2009 11:56:44 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id n1JJuD1m012151;
	Fri, 20 Feb 2009 03:56:39 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KFB00L09WQCYZ00@nwk-avmta-2.sfbay.sun.com>; Thu,
 19 Feb 2009 11:56:36 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KFB00J2IWQBY030@nwk-avmta-2.sfbay.sun.com>; Thu,
 19 Feb 2009 11:56:36 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id n1JJuSXX026879; Thu,
 19 Feb 2009 14:56:28 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id n1JJuSe4026876; Thu,
 19 Feb 2009 14:56:28 -0500 (EST)
Date: Thu, 19 Feb 2009 14:56:28 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: Issues for 2008/055
In-reply-to: <BF0EFF92-F807-4FD5-BD69-6FAF60A3DDC5@sun.com>
To: Nicolas Droux <droux@sun.com>
Cc: Sebastien Roy <Sebastien.Roy@sun.com>, PSARC-ext@sun.com
Message-id: <18845.47468.591337.751709@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <A4CDF502-BA36-43E8-A187-BB7B67488379@sun.com>
 <18761.26709.246151.818319@gargle.gargle.HOWL>
 <FEF6920A-4380-46B8-96C1-2D9E471C5C12@Sun.COM>
 <18762.31120.53165.970797@gargle.gargle.HOWL>
 <1491995A-D70F-4856-9627-36EC11B0B8CF@sun.com>
 <18770.22960.710148.605664@gargle.gargle.HOWL>
 <F9739781-85EB-491A-8D6A-B07ABD7F09DC@sun.com>
 <18837.52066.852097.442555@gargle.gargle.HOWL>
 <1234559753.22408.108.camel@strat>
 <BF0EFF92-F807-4FD5-BD69-6FAF60A3DDC5@sun.com>
Status: RO
Content-Length: 802

Nicolas Droux writes:
> > If we can agree that such work is indeed in the realm of a future
> > project, then I'd suggest including advice to that effect to the
> > project's management in the opinion for this case.
> 
> I agree with that approach. For the short term we will look into  
> providing a mechanism that will allow subsystems like bridging and  
> layer-2 filtering to intercept packets without requiring the addition  
> of checks specific to these subsystems in the common mac TX and RX  
> data-paths.

Thanks; yes, that sounds like the right way forward to me.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From carlsonj@phorcys.east.sun.com Fri Feb 20 14:32:34 2009
Received: from phorcys.east.sun.com (phorcys.East.Sun.COM [129.148.174.143])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n1KMWXSx016765
	for <psarc-ext@sac.sfbay.sun.com>; Fri, 20 Feb 2009 14:32:33 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id n1KMWQBM001509
	for <psarc-ext@sac.sfbay.sun.com>; Fri, 20 Feb 2009 17:32:26 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id n1KMWQOv001506;
	Fri, 20 Feb 2009 17:32:26 -0500 (EST)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Message-ID: <18847.12154.544711.937059@gargle.gargle.HOWL>
Date: Fri, 20 Feb 2009 17:32:26 -0500
From: James Carlson <james.d.carlson@sun.com>
To: psarc-ext@sac.sfbay.sun.com
Subject: 2008/055 Solaris Bridging: call for vote
X-Mailer: VM 7.01 under Emacs 21.3.1
Status: RO
Content-Length: 808

I've somewhat optimistically placed updated materials into
"final.materials/" subdirectory in the case directory.

In here, you will find:

  bridging-arc-changes.txt  summary of changes since inception
  bridging-design.pdf	    updated detailed design (informative reference)
  bridging-security.txt	    updated security checklist
  bridging-spec.txt	    updated architectural document (normative)

On Wednesday, I'll poll the members during ARC business.  If you're
not ready to vote or would like a full commitment review instead, then
that would be the right time to let me know.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From carlsonj@phorcys.east.sun.com Mon Jun  8 08:46:30 2009
Received: from dm-east-02.east.sun.com (dm-east-02.East.Sun.COM [129.148.13.5])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n58FkUog013298
	for <psarc-ext@sac.sfbay.sun.com>; Mon, 8 Jun 2009 08:46:30 -0700 (PDT)
Received: from phorcys.east.sun.com (phorcys.East.Sun.COM [129.148.174.143])
	by dm-east-02.east.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id n58FkTv4064151
	for <psarc-ext@sac.sfbay.sun.com>; Mon, 8 Jun 2009 11:46:29 -0400 (EDT)
Received: from phorcys.east.sun.com (phorcys.local [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id n58FjJEE013336
	for <psarc-ext@sac.sfbay.sun.com>; Mon, 8 Jun 2009 11:45:19 -0400 (EDT)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id n58FjJZq013333;
	Mon, 8 Jun 2009 11:45:19 -0400 (EDT)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Message-ID: <18989.12815.129599.40085@gargle.gargle.HOWL>
Date: Mon, 8 Jun 2009 11:45:19 -0400
From: James Carlson <james.d.carlson@sun.com>
To: psarc-ext@sac.sfbay.sun.com
Subject: Opinion for review: 2008/055 Solaris Bridging
X-Mailer: VM 7.01 under Emacs 21.3.1
Status: RO
Content-Length: 6268

Please review the following opinion and submit comments by COB on
06/12/2009.  Note that the timer is set a bit short for personal
reasons.

Note also that the opinion reflects the materials as reviewed by the
ARC.  For those who've participated in the design review, some things
(particularly /dev/bridge/) have changed since this ARC review was
completed, and those changes will be the subject of a fast-track to be
filed shortly.



 sun
   microsystems              Systems Architecture Committee

_________________________________________________________________

Subject:       Solaris Bridging

Submitted by:  James Carlson

File:          PSARC/2008/055/opinion.ms

Date:          February 25th, 2009

Committee:     James  D.  Carlson,  Kais  Belgaied,  Richard
               Matthews, Sebastien Roy.

Product Approval Committee:

               Solaris PAC
               solaris-pac@sun.com

1.  Summary

This project provides Ethernet  bridging  functionality  for
Solaris.

2.  Decision & Precedence Information

The project is approved as specified in reference [1].

The project may be delivered in a Minor release  of  Solaris
or OpenSolaris.

3.  Interfaces

The project exports the following interfaces.

____________________________________________________________________________
|                           Interfaces Exported                            |
|_____________________|_______________________|____________________________|
|Interface            |  Classification       |  Comments                  |
|_____________________|_______________________|____________________________|
|dladm *-bridge       |  Committed            |  new subcommands           |
|field names          |  Committed            |  dladm show-bridge -o      |
|link properties      |  Committed            |  dladm set-linkprop        |
|show-link BRIDGE     |  Committed            |  new field                 |
|kstats               |  Volatile             |  Should be raised later    |
|/dev/bridge/         |  Committed            |  Observability node        |
|control ioctls       |  Project Private      |                            |
|/usr/lib/bridged     |  Project Private      |  Daemon executable         |
|svc:/network/bridge  |  Committed            |  SMF URI                   |
|config/*             |  Project Private      |  SMF properties            |
|_____________________|_______________________|____________________________|

PSARC/2008/055               Copyright 2009 Sun Microsystems

                           - 2 -

____________________________________________________________________________
|                           Interfaces Exported                            |
|_____________________|_______________________|____________________________|
|Interface            |  Classification       |  Comments                  |
|_____________________|_______________________|____________________________|
|bridge module        |  Project Private      |  Kernel bridging module    |
|/var/run/bridge_door/|  Project Private      |  Doors interface to daemons|
|librstp.so.1         |  Project Private      |  RSTP implementation       |
|mac, dls, dld        |  Consolidation Private|  Kernel APIs               |
|::dladm show-bridge  |  Volatile             |  mdb dcmd (debugging)      |
|_____________________|_______________________|____________________________|

4.  Opinion

This project was originally filed as a fast-track, but  then
derailed  for  regular  review due to the depth of the ques-
tions raised.  At inception, the project team was advised to
consult  with the Crossbow and IP Filtering teams to resolve
the connections between these projects.   On  completion  of
those  discussions, the ARC members were updated (see refer-
ence [2]), and a vote on the final materials was held during
ARC business.

4.1.  IP Filter

The project team discussed filtering and bridging at length.
There  are  essentially  two  ways  that layer two filtering
(L2F) can apply to bridges: it  can  apply  on  top  of  the
bridge,  so that the links seen by L2F are the same as those
seen by IP, or it can apply below the bridge,  so  that  the
links  seen by L2F are the same as the physical links on the
system.

The former is expedient, but the  latter  will  require  new
interfaces,  including  a  "bridge  forwarding" hook that is
analogous to the existing "IP forwarding" hook.   This  work
is left to a future project to define.

4.2.  Crossbow

The bridging project allows  Crossbow's  flows  and  virtual
interfaces  to  be  used  on  top  of bridges for control of
traffic sent and received by local endpoints, but  does  not
make  use  of Crossbow's classification functionality in the
bridge forwarding function.  The project teams agree that it
would  be  better if this sort of integration were possible,
but the required functionality for  bridge  forwarding  does
not  currently  exist  in  Crossbow,  and retrofitting later
would be a seemless operation for users.   Thus,  the  teams
agreed  that  this future work can continue in parallel, and
that bridging should  be  reworked  when  suitable  Crossbow
interfaces are designed.

PSARC/2008/055               Copyright 2009 Sun Microsystems

                           - 3 -

4.3.  Security

An ARC member noted several problems and  complexities  with
the  originally proposed security mechanism.  The design [3]
was updated to drive all configuration through the  existing
SMF/SCF  and  dladm/dlmgmtd  interfaces,  so the project now
relies exclusively on existing security mechanisms  and  the
issues raised at inception are no longer present.

5.  Minority Opinion(s)

None

6.  Advisory Information

None

7.  Appendices

7.1.  Appendix A: Technical Changes Required

None

7.2.  Appendix B: Technical Changes Advised

None

7.3.  Appendix C: Reference Material

Unless stated otherwise, path names are relative to the case
directory PSARC/2008/055.

1.   Bridging Architectural Specification
     File:  final.materials/bridging-spec.txt

2.   ARC Update Summary
     File:  final.materials/bridging-arc-changes.txt

3.   Bridging Design Document
     File:  final.materials/bridging-design.pdf

PSARC/2008/055               Copyright 2009 Sun Microsystems


From sac-owner Fri Jun 12 15:28:32 2009
Received: from dm-east-01.east.sun.com (dm-east-01.East.Sun.COM [129.148.9.192])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n5CMSWeS005754
	for <sac-review@sac.sfbay.sun.com>; Fri, 12 Jun 2009 15:28:32 -0700 (PDT)
Received: from phorcys.east.sun.com (phorcys.East.Sun.COM [129.148.174.143])
	by dm-east-01.east.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id n5CMSVgJ060397
	for <sac-review@sac.sfbay.sun.com>; Fri, 12 Jun 2009 18:28:31 -0400 (EDT)
Received: from phorcys.east.sun.com (phorcys.local [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id n5CMRGx3028683
	for <sac-review@sac.sfbay.sun.com>; Fri, 12 Jun 2009 18:27:16 -0400 (EDT)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id n5CMRGae028680;
	Fri, 12 Jun 2009 18:27:16 -0400 (EDT)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Message-ID: <18994.54852.206741.284791@gargle.gargle.HOWL>
Date: Fri, 12 Jun 2009 18:27:16 -0400
From: James Carlson <james.d.carlson@sun.com>
To: sac-review@sac.sfbay.sun.com
Subject: opinion for review: PSARC 2008/055 Solaris Bridging
X-Mailer: VM 7.01 under Emacs 21.3.1
Status: RO
Content-Length: 6020

Please review the following opinion and submit comments by COB on
06/18/2009.  Note that the timer is set a bit short for personal
reasons.


 sun
   microsystems              Systems Architecture Committee

_________________________________________________________________

Subject:       Solaris Bridging

Submitted by:  James Carlson

File:          PSARC/2008/055/opinion.ms

Date:          February 25th, 2009

Committee:     James  D.  Carlson,  Kais  Belgaied,  Richard
               Matthews, Sebastien Roy.

Product Approval Committee:

               Solaris PAC
               solaris-pac@sun.com

1.  Summary

This project provides Ethernet  bridging  functionality  for
Solaris.

2.  Decision & Precedence Information

The project is approved as specified in reference [1].

The project may be delivered in a Minor release  of  Solaris
or OpenSolaris.

3.  Interfaces

The project exports the following interfaces.

____________________________________________________________________________
|                           Interfaces Exported                            |
|_____________________|_______________________|____________________________|
|Interface            |  Classification       |  Comments                  |
|_____________________|_______________________|____________________________|
|dladm *-bridge       |  Committed            |  new subcommands           |
|field names          |  Committed            |  dladm show-bridge -o      |
|link properties      |  Committed            |  dladm set-linkprop        |
|show-link BRIDGE     |  Committed            |  new field                 |
|kstats               |  Volatile             |  Should be raised later    |
|/dev/bridge/         |  Committed            |  Observability node        |
|control ioctls       |  Project Private      |                            |
|/usr/lib/bridged     |  Project Private      |  Daemon executable         |
|svc:/network/bridge  |  Committed            |  SMF URI                   |
|config/*             |  Project Private      |  SMF properties            |
|_____________________|_______________________|____________________________|

PSARC/2008/055               Copyright 2009 Sun Microsystems

                           - 2 -

____________________________________________________________________________
|                           Interfaces Exported                            |
|_____________________|_______________________|____________________________|
|Interface            |  Classification       |  Comments                  |
|_____________________|_______________________|____________________________|
|bridge module        |  Project Private      |  Kernel bridging module    |
|/var/run/bridge_door/|  Project Private      |  Doors interface to daemons|
|librstp.so.1         |  Project Private      |  RSTP implementation       |
|mac, dls, dld        |  Consolidation Private|  Kernel APIs               |
|::dladm show-bridge  |  Volatile             |  mdb dcmd (debugging)      |
|_____________________|_______________________|____________________________|

4.  Opinion

This project was originally filed as a fast-track, but  then
derailed  for  regular  review due to the depth of the ques-
tions raised.  At inception, the project team was advised to
consult  with the Crossbow and IP Filtering teams to resolve
the connections between these projects.   On  completion  of
those  discussions, the ARC members were updated (see refer-
ence [2]), and a vote on the final materials was held during
ARC business.

4.1.  IP Filter

The project team discussed filtering and bridging at length.
There  are  essentially  two  ways  that layer two filtering
(L2F) can apply to bridges: it  can  apply  on  top  of  the
bridge,  so that the links seen by L2F are the same as those
seen by IP, or it can apply below the bridge,  so  that  the
links  seen by L2F are the same as the physical links on the
system.

The former is expedient, but the  latter  will  require  new
interfaces,  including  a  "bridge  forwarding" hook that is
analogous to the existing "IP forwarding" hook.   This  work
is left to a future project to define.

4.2.  Crossbow

The bridging project allows  Crossbow's  flows  and  virtual
interfaces  to  be  used  on  top  of bridges for control of
traffic sent and received by local endpoints, but  does  not
make  use  of Crossbow's classification functionality in the
bridge forwarding function.  The project teams agree that it
would  be  better if this sort of integration were possible,
but the required functionality for  bridge  forwarding  does
not  currently  exist in Crossbow, and retrofitting bridging
to use new Crossbow interfaces at some future date would  be
a seamless operation for users.  Thus, the teams agreed that
this future work can continue in parallel, and that bridging
should  be  reworked  when  suitable Crossbow interfaces are
designed.

PSARC/2008/055               Copyright 2009 Sun Microsystems

                           - 3 -

4.3.  Security

An ARC member noted several problems and  complexities  with
the  originally proposed security mechanism.  The design [3]
was updated to drive all configuration through the  existing
SMF/SCF  and  dladm/dlmgmtd  interfaces,  so the project now
relies exclusively on existing security mechanisms  and  the
issues raised at inception are no longer present.

5.  Minority Opinion(s)

None

6.  Advisory Information

None

7.  Appendices

7.1.  Appendix A: Technical Changes Required

None

7.2.  Appendix B: Technical Changes Advised

None

7.3.  Appendix C: Reference Material

Unless stated otherwise, path names are relative to the case
directory PSARC/2008/055.

1.   Bridging Architectural Specification
     File:  final.materials/bridging-spec.txt

2.   ARC Update Summary
     File:  final.materials/bridging-arc-changes.txt

3.   Bridging Design Document
     File:  final.materials/bridging-design.pdf

PSARC/2008/055               Copyright 2009 Sun Microsystems


From sac-owner Fri Jun 19 05:46:53 2009
Received: from dm-east-01.east.sun.com (dm-east-01.East.Sun.COM [129.148.9.192])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n5JCkqtF004149
	for <sac-opinion@sac.sfbay.sun.com>; Fri, 19 Jun 2009 05:46:53 -0700 (PDT)
Received: from phorcys.east.sun.com (phorcys.East.Sun.COM [129.148.174.143])
	by dm-east-01.east.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id n5JCkqwk061639;
	Fri, 19 Jun 2009 08:46:52 -0400 (EDT)
Received: from phorcys.east.sun.com (phorcys.local [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id n5JCjVAl006772;
	Fri, 19 Jun 2009 08:45:31 -0400 (EDT)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id n5JCjV1o006769;
	Fri, 19 Jun 2009 08:45:31 -0400 (EDT)
MIME-Version: 1.0
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
Message-ID: <19003.34923.904002.110248@gargle.gargle.HOWL>
Date: Fri, 19 Jun 2009 08:45:31 -0400
From: James Carlson <james.d.carlson@sun.com>
To: sac-opinion@sac.sfbay.sun.com
cc: solaris-pac-opinion@sun.com
Subject: PSARC Opinion: 2008/055 Solaris Bridging
X-Mailer: VM 7.01 under Emacs 21.3.1
Status: RO
Content-Length: 5879


 sun
   microsystems              Systems Architecture Committee

_________________________________________________________________

Subject:       Solaris Bridging

Submitted by:  James Carlson

File:          PSARC/2008/055/opinion.ms

Date:          February 25th, 2009

Committee:     James  D.  Carlson,  Kais  Belgaied,  Richard
               Matthews, Sebastien Roy.

Product Approval Committee:

               Solaris PAC
               solaris-pac@sun.com

1.  Summary

This project provides Ethernet  bridging  functionality  for
Solaris.

2.  Decision & Precedence Information

The project is approved as specified in reference [1].

The project may be delivered in a Minor release  of  Solaris
or OpenSolaris.

3.  Interfaces

The project exports the following interfaces.

____________________________________________________________________________
|                           Interfaces Exported                            |
|_____________________|_______________________|____________________________|
|Interface            |  Classification       |  Comments                  |
|_____________________|_______________________|____________________________|
|dladm *-bridge       |  Committed            |  new subcommands           |
|field names          |  Committed            |  dladm show-bridge -o      |
|link properties      |  Committed            |  dladm set-linkprop        |
|show-link BRIDGE     |  Committed            |  new field                 |
|kstats               |  Volatile             |  Should be raised later    |
|/dev/bridge/         |  Committed            |  Observability node        |
|control ioctls       |  Project Private      |                            |
|/usr/lib/bridged     |  Project Private      |  Daemon executable         |
|svc:/network/bridge  |  Committed            |  SMF URI                   |
|config/*             |  Project Private      |  SMF properties            |
|_____________________|_______________________|____________________________|

PSARC/2008/055               Copyright 2009 Sun Microsystems

                           - 2 -

____________________________________________________________________________
|                           Interfaces Exported                            |
|_____________________|_______________________|____________________________|
|Interface            |  Classification       |  Comments                  |
|_____________________|_______________________|____________________________|
|bridge module        |  Project Private      |  Kernel bridging module    |
|/var/run/bridge_door/|  Project Private      |  Doors interface to daemons|
|librstp.so.1         |  Project Private      |  RSTP implementation       |
|mac, dls, dld        |  Consolidation Private|  Kernel APIs               |
|::dladm show-bridge  |  Volatile             |  mdb dcmd (debugging)      |
|_____________________|_______________________|____________________________|

4.  Opinion

This project was originally filed as a fast-track, but  then
derailed  for  regular  review due to the depth of the ques-
tions raised.  At inception, the project team was advised to
consult  with the Crossbow and IP Filtering teams to resolve
the connections between these projects.   On  completion  of
those  discussions, the ARC members were updated (see refer-
ence [2]), and a vote on the final materials was held during
ARC business.

4.1.  IP Filter

The project team discussed filtering and bridging at length.
There  are  essentially  two  ways  that layer two filtering
(L2F) can apply to bridges: it  can  apply  on  top  of  the
bridge,  so that the links seen by L2F are the same as those
seen by IP, or it can apply below the bridge,  so  that  the
links  seen by L2F are the same as the physical links on the
system.

The former is expedient, but the  latter  will  require  new
interfaces,  including  a  "bridge  forwarding" hook that is
analogous to the existing "IP forwarding" hook.   This  work
is left to a future project to define.

4.2.  Crossbow

The bridging project allows  Crossbow's  flows  and  virtual
interfaces  to  be  used  on  top  of bridges for control of
traffic sent and received by local endpoints, but  does  not
make  use  of Crossbow's classification functionality in the
bridge forwarding function.  The project teams agree that it
would  be  better if this sort of integration were possible,
but the required functionality for  bridge  forwarding  does
not  currently  exist in Crossbow, and retrofitting bridging
to use new Crossbow interfaces at some future date would  be
a seamless operation for users.  Thus, the teams agreed that
this future work can continue in parallel, and that bridging
should  be  reworked  when  suitable Crossbow interfaces are
designed.

PSARC/2008/055               Copyright 2009 Sun Microsystems

                           - 3 -

4.3.  Security

An ARC member noted several problems and  complexities  with
the  originally proposed security mechanism.  The design [3]
was updated to drive all configuration through the  existing
SMF/SCF  and  dladm/dlmgmtd  interfaces,  so the project now
relies exclusively on existing security mechanisms  and  the
issues raised at inception are no longer present.

5.  Minority Opinion(s)

None

6.  Advisory Information

None

7.  Appendices

7.1.  Appendix A: Technical Changes Required

None

7.2.  Appendix B: Technical Changes Advised

None

7.3.  Appendix C: Reference Material

Unless stated otherwise, path names are relative to the case
directory PSARC/2008/055.

1.   Bridging Architectural Specification
     File:  final.materials/bridging-spec.txt

2.   ARC Update Summary
     File:  final.materials/bridging-arc-changes.txt

3.   Bridging Design Document
     File:  final.materials/bridging-design.pdf

PSARC/2008/055               Copyright 2009 Sun Microsystems


