From sacadmin Tue Apr  8 20:55:16 2008
Received: from sac.sfbay.sun.com (localhost [127.0.0.1])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m393tG5k024021;
	Tue, 8 Apr 2008 20:55:16 -0700 (PDT)
Received: (from dr146992@localhost)
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8/Submit) id m393tG8p024017;
	Tue, 8 Apr 2008 20:55:16 -0700 (PDT)
Date: Tue, 8 Apr 2008 20:55:16 -0700 (PDT)
From: Darren Reed <dr146992@sac.sfbay.sun.com>
Message-Id: <200804090355.m393tG8p024017@sac.sfbay.sun.com>
To: PSARC-record@sac.sfbay.sun.com
Subject: Packet interception for the MAC layer [PSARC/2008/249 FastTrack timeout 04/15/2008]
Status: RO
Content-Length: 568


Template Version: @(#)sac_nextcase 1.64 07/13/07 SMI
This information is Copyright 2008 Sun Microsystems
1. Introduction
    1.1. Project/Component Working Name:
	 Packet interception for the MAC layer
    1.2. Name of Document Author/Supplier:
	 Author:  Zhijun Fu
    1.3  Date of This Document:
	08 April, 2008
4. Technical Description
    See the case directory for more detail

6. Resources and Schedule
    6.4. Steering Committee requested information
   	6.4.1. Consolidation C-team Name:
		ON
    6.5. ARC review type: FastTrack
    6.6. ARC Exposure: open


From Darren.Reed@sun.com Tue Apr  8 22:55:47 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m395tkLM026432
	for <psarc-ext@sac.sfbay.Sun.COM>; Tue, 8 Apr 2008 22:55:47 -0700 (PDT)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id m395tjiJ019202
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Wed, 9 Apr 2008 13:55:46 +0800 (SGT)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ100305N4WDQ00@brm-avmta-1.central.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Tue, 08 Apr 2008 23:55:44 -0600 (MDT)
Received: from sineb-mail-2.sun.com ([192.18.19.7])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ1002OCN4ND600@brm-avmta-1.central.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Tue,
 08 Apr 2008 23:55:36 -0600 (MDT)
Received: from fe-apac-05.sun.com
 (fe-apac-05.sun.com [192.18.19.176] (may be forged))
	by sineb-mail-2.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m395u55j027602	for
 <psarc-ext@sun.com>; Wed, 09 Apr 2008 05:56:05 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JZ100301N257Y00@mail-apac.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Wed,
 09 Apr 2008 13:55:28 +0800 (SGT)
Received: from [129.146.106.55] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JZ10042LN4FPGBC@mail-apac.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Wed, 09 Apr 2008 13:55:28 +0800 (SGT)
Date: Tue, 08 Apr 2008 22:55:33 -0700
From: Darren Reed <Darren.Reed@sun.com>
Subject: PSARC/2008/249 Packet interception for the MAC layer
Sender: Darren.Reed@sun.com
To: PSARC-EXT <psarc-ext@sun.com>
Cc: Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <47FC5A55.2050405@Sun.COM>
MIME-version: 1.0
Content-type: multipart/mixed; boundary="Boundary_(ID_l0L5VEgCrNa5P/qsbzo7vg)"
X-Accept-Language: en-au, en
X-PMX-Version: 5.4.1.325704
User-Agent: Mozilla/5.0 (X11; U; SunOS i86pc; en-US; rv:1.7) Gecko/20060120
Status: RO
Content-Length: 9422

This is a multi-part message in MIME format.

--Boundary_(ID_l0L5VEgCrNa5P/qsbzo7vg)
Content-type: text/plain; format=flowed; charset=us-ascii
Content-transfer-encoding: 7BIT

I'm submitting the attached spec as the proposal for the Layer 2 
Filtering project
on behalf of Zhijun Fu.

Darren

--Boundary_(ID_l0L5VEgCrNa5P/qsbzo7vg)
Content-type: text/plain; name=l2filter_spec.txt
Content-transfer-encoding: 7BIT
Content-disposition: inline; filename=l2filter_spec.txt

Abstract
========
This case will extend PSARC/2005/334, by adding the ability to intercept
packets in MAC layer using the PFHooks infrastructure.

This case only makes one change, an addition, to the interfaces
that were committed to by PSARC/2008/219 (see "new hook event"
below for more details.)

Release Biding
--------------
This case seeks for a patch binding.

Background
==========
The PFHooks project, PSARC/2005/334, provide the ability to intercept packets
in IP layer by adding Hooks into network stack. 

Since its integration, there has been customer requirements for the ability
to intercept packets in MAC layer, also the ability is needed in order to
enforce security for xVM/Zone.

Introduction
============
Boundaries
----------
This case is confined to providing the ability to intercept inbound/outbound
packets from/to a network. Though this case will make it possible to adding 
new hooks to intercept inter-domain traffic for xVM, this project won't
deliver the hooks to implement this.

Goals
-----
This case seeks to meet the following goals:
* provide the hooks in MAC layer that allows consumers to register on to 
  intercept packets;

* provide the netinfo interface for ethernet that gives consumers access to 
  interface information, and the ability to inject or emit packets directly;

* modify ipfilter to provide the ability to filter layer 2 packets as ethernet
  packets, and also filter them as IP packets and do IP NAT if required.

Out of scope
------------
This case is concerned solely on intercepting inbound/outbound packets, thus
the following is considered out of scope:

* provide the ability to intercept inter-domain xVM packets.

Details
=======
This case would like to propose adding a new family of hooks for "ethernet".
 will make it possible to intercept packets in MAC layer.

netinfo & hook callback 
-----------------------
The hooks provided by ethernet will generate events for NH_PHYSICAL_IN and
NH_PHYSICAL_OUT, using the same interface as IPv4 and IPv6 do in 
PSARC/2005/334.

The following functions will be supported through the netinfo framework:
net_getifname()
net_phylookup()
net_phygetnext()
net_getlifaddr()
net_inject()

All of the other functions in the netinfo() framework will return a value
indicating that they are unsupported. The return values for the above 
functions only have meaning with the scope of ethernet - it is not correct
to use a value returned by net_getifname() using the ethernet net_data_t
handle with net_phylookup() for IPv4.

The callback for NH_PHYSICAL_IN and NH_PHYSICAL_OUT will receive a
pointer to a hook_packet_event_t structure that has the following
fields filled out:

hpe_ifp - 0 for NH_PHYSICAL_OUT, otherwise a value indicating which
          interface the NH_PHYSICAL_IN event is associated with;
hpe_ofp - 0 for NH_PHYSICAL_IN, otherwise a value indicating which
          interface the NH_PHYSICAL_OUT event is associated with;
hpe_hdr - points to the start of the ethernet header
hpe_mb  - points to the start of the mblk_t that holds hpe_hdr;
hpe_mp  - points to the mblk_t that is the start of the packet.

new hook event
--------------
As Clearview UV (PSARC/2006/499, PSARC/2007/527, PSARC/2008/002) introduces
the ability to rename a data link, we need to capture this event in order to
update ipfilter rules accrodingly. Thus we propose an extension to
PSARC/2008/219 by adding a new hook event NE_NAME_CHANGE to nic_event_t
to indicate the rename link event.

typedef enum nic_event {
         NE_PLUMB = 1,
         NE_UNPLUMB,
         NE_UP,
         NE_DOWN,
         NE_ADDRESS_CHANGE,
+        NE_NAME_CHANGE
} nic_event_t;

ipfilter changes
----------------
Users can use ipf(1M) to add ethernet filtering rules in addition to IP 
filtering rules, the ethernet filtering rules are marked with "family ether".
Unlike IPv6, no special command line switch is required to load ethernet 
rules. And by default, ethernet rules should be put in /etc/ipf/ipf.conf.

The layer 2 filtering functionality is disabled by default, to enable it,
add the following line to the top of ipf.conf:

set intercept_layer2 true;

Also, to provide the ability to process IP filtering & IP NAT rules in
MAC layer, two more keywords "ip-head", "ip-nat" are added.
If a packet matches an ethernet filtering rule which specifies "ip-head"
keyword, the packet will go to the corresponding IP filtering group
to be processed before it is passed up. Similarly, if a packet matches
an ethernet rule which specifies "ip-nat", the packet will be passed to
IP NAT rules to be NAT'ed before passed up.

To distinguish IP filter/NAT rules intended to be processed in layer 2
from the rest of ipfilter rules, an additional keyword "layer2" is added.
Those ipfilter rules to be processed in layer 2 are marked with "layer2",
so these rules won't be processed again when packets goes up to IP.

Below is an example:

(1) pass in on nge1 family ether from 0:14:4f:8d:ae:23 to any ip-head 10 ip-nat
(2) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
    keep state group 10 layer2 
(3) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32
(4) rdr nge1 from 10.10.10.20/32 to 10.10.10.10/32 port = 7777 -> 10.10.10.10
    port 8888 tcp layer2

In the above case, if a packet matches rule (1) in MAC layer, rule (2) and 
rule (4) will be processed in MAC layer for this packet, and when it goes
up to IP, rule (3) will be processed.

Also, ipmon has been updated to print out log records with ethernet
information but the output of this command is volatile.

link name mapping 
-----------------
For layer 2 rules, as well as ipfilter rules to be processed in layer 2,
we need to use mac_impl_t pointer as an interface indentifier in kernel.

Since Clearview UV integration, administrators need to use link name
for data link related operations.
As link name is used to specify the interface when adding an ipfilter rules,
we need to map the link name to a mac_impl_t pointer.
Also, ipfilter needs to map a mac_impl_t pointer to link name, in order to
generate correct logging message. Clearview UV provides dls_mgmt_get_linkid()
and dls_mgmt_get_linkinfo() to translate between link name and link id.
But in this case, the logging code is in data path thus will happen in
interrupt context, thus the above routines cannot be used because they
use door call to get the information.
Thus we propose to add a link name <-> link id hash table in dls, and provide
the following routines to translate between link name and mac name.
And ethernet netinfo will use these two routines to implement mapping between
link name and mac_impl_t pointer.

+----------------------------------------------------+
| Interface                         | Classification |
|----------------------------------------------------|
| dls_devnet_mac2link(const char *, | private        |
|     char *, const size_t);        |                |
| dls_devnet_link2mac(const char *, | private        |
|     char *, const size_t);        |                |
+----------------------------------------------------+
Table: Fuctions for link name/mac name mapping

Doc changes
-----------
* ipf(4)
+     l2filter-rule = eaction in-out [ eoptions ] ether .
+     eaction ::= "pass" | "block" | "log" | "count" | auth .
+     eoptions ::= [ "log" ] [ "quick" ] [ "on" interface-name ] .
+     ether ::= "family ether" [ "type" ether-type ] { "all" | efromto }
+               [ "vlan" decnumber ] [ "ip-head" decnumber ] [ "ip-nat" ] .
+     ether-type ::= hexdigit [ hexdigit [ hexdigit [ hexdigit] ] ] ] .
+     efromto ::= "from" eaddr "to" eaddr .
+     eaddr ::= "any" | ethaddr [ "/" decnumber ] .
+     ethaddr ::= eth-num ":" eth-num ":" eth-num ":" eth-num ":" eth-num ":"
+                 eth-num .
+     eth-num ::= hexdigit [ hexdigit ] .

      filter-rule = [ insert ] action in-out [ options ] [ tos ] [ ttl ]
-	 [ proto ] ip [ group ] .
+	 [ proto ] ip [ group ] [ "layer2" ] .

+  Layer 2 filtering
+     By default, Layer 2 filtering will not be enabled.
+
+     To enable layer 2 filtering, you must add the following line to ipf.conf
+     file:
+       
+       set intercept_layer2 true;
+
+     This line must be placed before any ipfilter or layer 2 filter rules
+     in this file.

* ipnat(4)
       map ::= mapit ifname ipmask "->" dstipmask [ mapport | mapproxy ] \
-              mapoptions.
+              mapoptions [ "layer2 " ].
-       map ::= mapit ifname fromto "->" dstipmask [ mapport ] mapoptions.
+       map ::= mapit ifname fromto "->" dstipmask [ mapport ] mapoptions
+            [ "layer2" ].
       mapblock ::= "map-block" ifname ipmask "->" ipmask [ ports ] \
-                    mapoptions.
+                    mapoptions [ "layer2 " ].
       redir ::= "rdr" ifname ipmask dport "->" ip [ "," ip ] rdrport \
-                 rdroptions .
+                 rdroptions [ "layer2 " ].

--Boundary_(ID_l0L5VEgCrNa5P/qsbzo7vg)--

From Darren.Moffat@sun.com Wed Apr  9 02:44:26 2008
Received: from sunmail2sca.sfbay.sun.com (sunmail2sca.SFBay.Sun.COM [129.145.155.234])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m399iQVo004752
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 9 Apr 2008 02:44:26 -0700 (PDT)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail2sca.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m399iQNC022462
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Wed, 9 Apr 2008 02:44:26 -0700 (PDT)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ100103XQ2H000@nwk-avmta-2.sfbay.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Wed, 09 Apr 2008 02:44:26 -0700 (PDT)
Received: from gmp-eb-inf-1.sun.com ([192.18.6.21])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ100EC5XQ1CXF0@nwk-avmta-2.sfbay.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Wed,
 09 Apr 2008 02:44:25 -0700 (PDT)
Received: from fe-emea-10.sun.com (gmp-eb-lb-2-fe1.eu.sun.com [192.18.6.10])
	by gmp-eb-inf-1.sun.com (8.13.7+Sun/8.12.9) with ESMTP id m399iOcg013246	for
 <psarc-ext@sun.com>; Wed, 09 Apr 2008 09:44:24 +0000 (GMT)
Received: from conversion-daemon.fe-emea-10.sun.com by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0JZ100101VZ9CT00@fe-emea-10.sun.com>
 (original mail from Darren.Moffat@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Wed,
 09 Apr 2008 10:44:24 +0100 (BST)
Received: from [129.156.173.199] by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0JZ100A6FXQ0S530@fe-emea-10.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Wed, 09 Apr 2008 10:44:24 +0100 (BST)
Date: Wed, 09 Apr 2008 10:44:24 +0100
From: Darren J Moffat <Darren.Moffat@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <47FC5A55.2050405@Sun.COM>
Sender: Darren.Moffat@sun.com
To: Darren Reed <Darren.Reed@sun.com>
Cc: PSARC-EXT <PSARC-ext@sun.com>, Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <47FC8FF8.8010301@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
User-Agent: Thunderbird 2.0.0.9 (X11/20080225)
Status: RO
Content-Length: 217

Why is layer2 filtering disabled by default ?

What happens if filtering rules that depend on layer2 processing are 
added (such as the ones in the example) but intercept_layer2 hasn't been 
set ?

--
Darren J Moffat

From Zhijun.Fu@sun.com Wed Apr  9 03:35:33 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m39AZXj2006591
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 9 Apr 2008 03:35:33 -0700 (PDT)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m39AZWUr044120
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Wed, 9 Apr 2008 04:35:33 -0600 (MDT)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ200301038IH00@nwk-avmta-2.sfbay.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Wed, 09 Apr 2008 03:35:32 -0700 (PDT)
Received: from sineb-mail-1.sun.com ([192.18.19.6])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ2001LT032IG20@nwk-avmta-2.sfbay.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Wed,
 09 Apr 2008 03:35:27 -0700 (PDT)
Received: from fe-apac-05.sun.com
 (fe-apac-05.sun.com [192.18.19.176] (may be forged))
	by sineb-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m39AZcct024125	for
 <PSARC-ext@sun.com>; Wed, 09 Apr 2008 10:35:38 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JZ200F0102Q1O00@mail-apac.sun.com> (original mail from Zhijun.Fu@Sun.COM)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Wed,
 09 Apr 2008 18:35:18 +0800 (SGT)
Received: from [129.158.215.37] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JZ2004WF02UPGAE@mail-apac.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Wed, 09 Apr 2008 18:35:18 +0800 (SGT)
Date: Wed, 09 Apr 2008 18:34:29 +0800
From: Zhijun Fu <Zhijun.Fu@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <47FC8FF8.8010301@Sun.COM>
Sender: Zhijun.Fu@sun.com
To: Darren J Moffat <Darren.Moffat@sun.com>
Cc: Darren Reed <Darren.Reed@sun.com>, PSARC-EXT <PSARC-ext@sun.com>
Reply-to: Zhijun.Fu@sun.com
Message-id: <47FC9BB5.7050505@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM> <47FC8FF8.8010301@Sun.COM>
User-Agent: Thunderbird 2.0.0.6 (X11/20071119)
Status: RO
Content-Length: 942

Darren J Moffat wrote:
> Why is layer2 filtering disabled by default ?
This is to keep system behavior consistent with before by default.
>
> What happens if filtering rules that depend on layer2 processing are 
> added (such as the ones in the example) but intercept_layer2 hasn't 
> been set ?
In this case, the addition of these rules will fail.  Such as in the 
example, the addition of rule (1), (2), (4) will fail if 
intercept_layer2 is not set.
Layer 2 rules and ipfilter rules that depend on layer2 processing will 
only be able to be added after the intercept_layer2 is set, and these 
rules will be flushed when the flag is unset.

Thanks,
Zhijun
>
> -- 
> Darren J Moffat


-- 
#mdb -K
[0]> eri.prc.sun.com::walk staff s|::print staff_t s_email|
::grep .== Zhijun.Fu@Sun.COM|::eval <s=K|::print staff_t
Zhijun.Fu@Sun.COM, x84349
Network Virtualization & Performance Team,
Solaris Core Operating Systems
Since Jul 10,2006
[0]> :c


From Darren.Moffat@Sun.COM Wed Apr  9 03:42:08 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m39Ag7AN006623
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 9 Apr 2008 03:42:07 -0700 (PDT)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id m39Ag6YT026947
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Wed, 9 Apr 2008 11:42:06 +0100 (BST)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ2003050E5RK00@nwk-avmta-2.sfbay.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@Sun.COM); Wed, 09 Apr 2008 03:42:05 -0700 (PDT)
Received: from gmp-eb-inf-1.sun.com ([192.18.6.21])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ2001ZS0E4IG20@nwk-avmta-2.sfbay.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@Sun.COM); Wed,
 09 Apr 2008 03:42:05 -0700 (PDT)
Received: from fe-emea-10.sun.com (gmp-eb-lb-2-fe1.eu.sun.com [192.18.6.10])
	by gmp-eb-inf-1.sun.com (8.13.7+Sun/8.12.9) with ESMTP id m39Ag4VG022054	for
 <PSARC-ext@Sun.COM>; Wed, 09 Apr 2008 10:42:04 +0000 (GMT)
Received: from conversion-daemon.fe-emea-10.sun.com by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0JZ100701ZV87100@fe-emea-10.sun.com>
 (original mail from Darren.Moffat@Sun.COM)
 for PSARC-ext@Sun.COM (ORCPT PSARC-ext@Sun.COM); Wed,
 09 Apr 2008 11:42:04 +0100 (BST)
Received: from [129.156.173.199] by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0JZ200D620DHVI00@fe-emea-10.sun.com> for PSARC-ext@Sun.COM
 (ORCPT PSARC-ext@Sun.COM); Wed, 09 Apr 2008 11:41:42 +0100 (BST)
Date: Wed, 09 Apr 2008 11:41:41 +0100
From: Darren J Moffat <Darren.Moffat@Sun.COM>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <47FC9BB5.7050505@Sun.COM>
Sender: Darren.Moffat@Sun.COM
To: Zhijun.Fu@Sun.COM
Cc: Darren Reed <Darren.Reed@Sun.COM>, PSARC-EXT <PSARC-ext@Sun.COM>
Message-id: <47FC9D65.9010405@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM> <47FC8FF8.8010301@Sun.COM>
 <47FC9BB5.7050505@Sun.COM>
User-Agent: Thunderbird 2.0.0.9 (X11/20080225)
Status: RO
Content-Length: 539

Zhijun Fu wrote:
> Darren J Moffat wrote:
>> Why is layer2 filtering disabled by default ?
> This is to keep system behavior consistent with before by default.

What would change in behaviour if layer2 filtering was on by default ?

So far I don't actually see any reason this needs to be configurable at all.

Are there existing rules that don't mention layer2 that would cause 
different filtering decisions if layer2 filtering was always on ?

Is there a performance impact from layer2 filtering always being on ?


-- 
Darren J Moffat

From Zhijun.Fu@sun.com Wed Apr  9 05:03:36 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m39C3ZSH008221
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 9 Apr 2008 05:03:35 -0700 (PDT)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id m39C3U25026222
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Wed, 9 Apr 2008 13:03:34 +0100 (BST)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ200H0J45VNE00@nwk-avmta-1.sfbay.Sun.COM> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@Sun.COM); Wed, 09 Apr 2008 05:03:31 -0700 (PDT)
Received: from sineb-mail-2.sun.com ([192.18.19.7])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ200ENU45ULBB0@nwk-avmta-1.sfbay.Sun.COM> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@Sun.COM); Wed,
 09 Apr 2008 05:03:31 -0700 (PDT)
Received: from fe-apac-05.sun.com
 (fe-apac-05.sun.com [192.18.19.176] (may be forged))
	by sineb-mail-2.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m39C40bG022695	for
 <PSARC-ext@Sun.COM>; Wed, 09 Apr 2008 12:04:00 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JZ20060140DAJ00@mail-apac.sun.com> (original mail from Zhijun.Fu@Sun.COM)
 for PSARC-ext@Sun.COM (ORCPT PSARC-ext@Sun.COM); Wed,
 09 Apr 2008 20:03:22 +0800 (SGT)
Received: from [129.158.215.37] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JZ2004XB45LPGPE@mail-apac.sun.com> for PSARC-ext@Sun.COM
 (ORCPT PSARC-ext@Sun.COM); Wed, 09 Apr 2008 20:03:22 +0800 (SGT)
Date: Wed, 09 Apr 2008 20:02:33 +0800
From: Zhijun Fu <Zhijun.Fu@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <47FC9D65.9010405@Sun.COM>
Sender: Zhijun.Fu@sun.com
To: Darren J Moffat <Darren.Moffat@sun.com>
Cc: Darren Reed <Darren.Reed@sun.com>, PSARC-EXT <PSARC-ext@sun.com>
Reply-to: Zhijun.Fu@sun.com
Message-id: <47FCB059.5090602@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM> <47FC8FF8.8010301@Sun.COM>
 <47FC9BB5.7050505@Sun.COM> <47FC9D65.9010405@Sun.COM>
User-Agent: Thunderbird 2.0.0.6 (X11/20071119)
Status: RO
Content-Length: 1242

Darren J Moffat wrote:
> Zhijun Fu wrote:
>> Darren J Moffat wrote:
>>> Why is layer2 filtering disabled by default ?
>> This is to keep system behavior consistent with before by default.
>
> What would change in behaviour if layer2 filtering was on by default ?
There will be performance impact, please see below
>
> So far I don't actually see any reason this needs to be configurable 
> at all.
>
> Are there existing rules that don't mention layer2 that would cause 
> different filtering decisions if layer2 filtering was always on ?
No,  the behavior of these rules will be the same whether or not layer2 
filtering is enabled.
>
> Is there a performance impact from layer2 filtering always being on ?
Yes.  There will be performance impact if layer2 filtering is always on, 
even if there are no layer2 rules configured,  because additional 
processing is needed when layer2 filtering is enabled.

This is similar to ipfilter, which is also disabled by default.

Thanks,
Zhijun

-- 
#mdb -K
[0]> eri.prc.sun.com::walk staff s|::print staff_t s_email|
::grep .== Zhijun.Fu@Sun.COM|::eval <s=K|::print staff_t
Zhijun.Fu@Sun.COM, x84349
Network Virtualization & Performance Team,
Solaris Core Operating Systems
Since Jul 10,2006
[0]> :c


From carlsonj@phorcys.east.sun.com Wed Apr  9 06:07:36 2008
Received: from sunmail3mpk.sfbay.sun.com (sunmail3mpk.SFBay.Sun.COM [129.146.11.52])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m39D7ak3009994
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 9 Apr 2008 06:07:36 -0700 (PDT)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail3mpk.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m39D7YmI018432;
	Wed, 9 Apr 2008 06:07:34 -0700 (PDT)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ20090L74LR300@brm-avmta-1.central.sun.com>; Wed,
 09 Apr 2008 07:07:33 -0600 (MDT)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ2004P174IV730@brm-avmta-1.central.sun.com>; Wed,
 09 Apr 2008 07:07:31 -0600 (MDT)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m39D7RUh009416; Wed,
 09 Apr 2008 09:07:27 -0400 (EDT)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m39D7RQJ009413; Wed,
 09 Apr 2008 09:07:27 -0400 (EDT)
Date: Wed, 09 Apr 2008 09:07:27 -0400
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <47FC5A55.2050405@Sun.COM>
To: Darren Reed <Darren.Reed@sun.com>
Cc: PSARC-EXT <psarc-ext@sun.com>, Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <18428.49039.503949.2406@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
Status: RO
Content-Length: 3013

Darren Reed writes:
> This case will extend PSARC/2005/334, by adding the ability to intercept
> packets in MAC layer using the PFHooks infrastructure.

Very minor nit: case title doesn't quite match the contents.  I was
excited to see the name, because I'm working on MAC layer interception
... until I read that it was just a PFHooks extension for layer 2.

> Users can use ipf(1M) to add ethernet filtering rules in addition to IP 
> filtering rules, the ethernet filtering rules are marked with "family ether".
> Unlike IPv6, no special command line switch is required to load ethernet 
> rules. And by default, ethernet rules should be put in /etc/ipf/ipf.conf.

That seems strange.

We currently have /etc/ipf/ipf.conf for IPv4 and the undocumented
/etc/ipf/ipf6.conf for IPv6.  Why wouldn't we have /etc/ipf/ipfl2.conf
(or some such) for L2-specific rules?

Or if "family ether" is a good way to do this, why wouldn't we have
"family inet" and "family inet6" and get rid of /etc/ipf/ipf6.conf?

What's the intended direction?

> The layer 2 filtering functionality is disabled by default, to enable it,
> add the following line to the top of ipf.conf:
> 
> set intercept_layer2 true;

If I have to say "family ether" in order to specify a rule that
filters L2 packets, why do I need to give this extra command?  Doesn't
the existence of at least one "family ether" rule mean that I intend
to filter L2 packets as well (and thus I want interception turned on)?

> To distinguish IP filter/NAT rules intended to be processed in layer 2
> from the rest of ipfilter rules, an additional keyword "layer2" is added.
> Those ipfilter rules to be processed in layer 2 are marked with "layer2",
> so these rules won't be processed again when packets goes up to IP.

If I have rules that have "layer2" set, then why do I need to specify
"ip-head" or "ip-nat" in the "family ether" filter?  Shouldn't any
rules with "layer2" set just _automatically_ match?

Or perhaps the question is this: why would I want to have rules
specified as "layer2", but then specifically avoid sending some
packets through those rules with "ip-head" or "ip-nat"?  If I did have
such a case, why wouldn't I set up a "family ether" rule that
specifies "quick" -- so that the rest of the "layer2"-tagged rules
aren't examined at all?  That (using "quick" instead for the reverse
sense) seems a lot clearer to me than "ip-head" or "ip-nat".

I suspect that most sane rule sets will start with something like
this:

	pass in on nge1 family ether all ip-head ip-nat

... so that the remaining "layer2" filter entries (who filters on
explicit MAC addresses?) won't be confusing.

(It seems to me that "family ether" and "layer2" are essentially the
same thing; they're in lieu of putting these rules in a separate
file.)

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Darren.Reed@sun.com Wed Apr  9 10:02:24 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m39H2O8m021538
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 9 Apr 2008 10:02:24 -0700 (PDT)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m39H2Ls3046882
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Wed, 9 Apr 2008 11:02:23 -0600 (MDT)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ200L1XHZW4A00@nwk-avmta-1.sfbay.Sun.COM> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Wed, 09 Apr 2008 10:02:20 -0700 (PDT)
Received: from sineb-mail-2.sun.com ([192.18.19.7])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ2008NDHZRWUD0@nwk-avmta-1.sfbay.Sun.COM> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Wed,
 09 Apr 2008 10:02:16 -0700 (PDT)
Received: from fe-apac-06.sun.com
 (fe-apac-06.sun.com [192.18.19.177] (may be forged))
	by sineb-mail-2.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m39H2jqf000518	for
 <psarc-ext@sun.com>; Wed, 09 Apr 2008 17:02:45 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JZ200701HO52200@mail-apac.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Thu,
 10 Apr 2008 01:01:42 +0800 (SGT)
Received: from [129.146.106.55] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JZ200BWEHYRRED3@mail-apac.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Thu, 10 Apr 2008 01:01:42 +0800 (SGT)
Date: Wed, 09 Apr 2008 10:02:11 -0700
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <18428.49039.503949.2406@gargle.gargle.HOWL>
Sender: Darren.Reed@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: PSARC-EXT <psarc-ext@sun.com>, Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <47FCF693.1040209@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-Accept-Language: en-au, en
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
 <18428.49039.503949.2406@gargle.gargle.HOWL>
User-Agent: Mozilla/5.0 (X11; U; SunOS i86pc; en-US; rv:1.7) Gecko/20060120
Status: RO
Content-Length: 1716

James Carlson wrote:

>Darren Reed writes:
>  
>
>>This case will extend PSARC/2005/334, by adding the ability to intercept
>>packets in MAC layer using the PFHooks infrastructure.
>>    
>>
>
>Very minor nit: case title doesn't quite match the contents.  I was
>excited to see the name, because I'm working on MAC layer interception
>... until I read that it was just a PFHooks extension for layer 2.
>
>  
>
>>Users can use ipf(1M) to add ethernet filtering rules in addition to IP 
>>filtering rules, the ethernet filtering rules are marked with "family ether".
>>Unlike IPv6, no special command line switch is required to load ethernet 
>>rules. And by default, ethernet rules should be put in /etc/ipf/ipf.conf.
>>    
>>
>
>That seems strange.
>
>We currently have /etc/ipf/ipf.conf for IPv4 and the undocumented
>/etc/ipf/ipf6.conf for IPv6.  Why wouldn't we have /etc/ipf/ipfl2.conf
>(or some such) for L2-specific rules?
>
>Or if "family ether" is a good way to do this, why wouldn't we have
>"family inet" and "family inet6" and get rid of /etc/ipf/ipf6.conf?
>
>What's the intended direction?
>  
>


In PSARC/2005/201, which was IPv6 for IPFilter, the direction from
PSARC was to move to a single configuration file for all of the filtering
statements - thus /etc/ipf/ipf6.conf was introduced as an obsolete
interface with the understanding that it would be subsumed in the
future by /etc/ipf/ipf.conf.  The background here is that the current
use of Ipv6 filtering outside of Solaris uses a separate file.  Thus it
seemed to not make any sense to introduce a new file that would also
be obsolete at introduction - more importantly, there is no prior history
in open source for a separate file.

Darren


From carlsonj@phorcys.east.sun.com Wed Apr  9 10:08:26 2008
Received: from sunmail3mpk.sfbay.sun.com (sunmail3mpk.SFBay.Sun.COM [129.146.11.52])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m39H8PR6022073
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 9 Apr 2008 10:08:26 -0700 (PDT)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail3mpk.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m39H8Ndh014901;
	Wed, 9 Apr 2008 10:08:23 -0700 (PDT)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ200G0XI9YGY00@nwk-avmta-2.sfbay.sun.com>; Wed,
 09 Apr 2008 10:08:22 -0700 (PDT)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ200BT4I9WINB0@nwk-avmta-2.sfbay.sun.com>; Wed,
 09 Apr 2008 10:08:20 -0700 (PDT)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m39H8JUW010534; Wed,
 09 Apr 2008 13:08:19 -0400 (EDT)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m39H8Jk9010531; Wed,
 09 Apr 2008 13:08:19 -0400 (EDT)
Date: Wed, 09 Apr 2008 13:08:19 -0400
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <47FCF693.1040209@Sun.COM>
To: Darren Reed <Darren.Reed@sun.com>
Cc: PSARC-EXT <psarc-ext@sun.com>, Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <18428.63491.785639.668640@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
 <18428.49039.503949.2406@gargle.gargle.HOWL> <47FCF693.1040209@Sun.COM>
Status: RO
Content-Length: 1620

Darren Reed writes:
> >Or if "family ether" is a good way to do this, why wouldn't we have
> >"family inet" and "family inet6" and get rid of /etc/ipf/ipf6.conf?
> >
> >What's the intended direction?
> >  
> >
> 
> 
> In PSARC/2005/201, which was IPv6 for IPFilter, the direction from
> PSARC was to move to a single configuration file for all of the filtering
> statements - thus /etc/ipf/ipf6.conf was introduced as an obsolete
> interface with the understanding that it would be subsumed in the
> future by /etc/ipf/ipf.conf.

Sure.  What's confusing me here is that we're not actually getting
that merge.  Instead, we're getting something new grafted onto
/etc/ipf/ipf.conf, while IPv6 remains an outpost in
/etc/ipf/ipf6.conf.

>  The background here is that the current
> use of Ipv6 filtering outside of Solaris uses a separate file.  Thus it
> seemed to not make any sense to introduce a new file that would also
> be obsolete at introduction - more importantly, there is no prior history
> in open source for a separate file.

OK ... so if I want to filter IPv6 packets using the new L2 mechanism,
do I put the IPv6 rules into /etc/ipf/ipf.conf alone or do the "family
ether" bits go into /etc/ipf/ipf.conf with the v6 "layer2"-tagged
rules in /etc/ipf/ipf6.conf?

(And, assuming you're not taking the other comments, do I then need
"ip6-head" and perhaps even "ip6-nat" as directives?)

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Zhijun.Fu@sun.com Wed Apr  9 19:49:07 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m3A2n6XY015880
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 9 Apr 2008 19:49:07 -0700 (PDT)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id m3A2n1Ea029069
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Thu, 10 Apr 2008 03:49:05 +0100 (BST)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ300G0B95RKB00@nwk-avmta-2.sfbay.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Wed, 09 Apr 2008 19:49:03 -0700 (PDT)
Received: from sineb-mail-1.sun.com ([192.18.19.6])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ3007FQ95QJSC0@nwk-avmta-2.sfbay.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Wed,
 09 Apr 2008 19:49:03 -0700 (PDT)
Received: from fe-apac-06.sun.com
 (fe-apac-06.sun.com [192.18.19.177] (may be forged))
	by sineb-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m3A2nErW023381	for
 <psarc-ext@sun.com>; Thu, 10 Apr 2008 02:49:14 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JZ300K018ZW9A00@mail-apac.sun.com> (original mail from Zhijun.Fu@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Thu,
 10 Apr 2008 10:48:29 +0800 (SGT)
Received: from [129.158.215.37] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JZ300BE794SRFAD@mail-apac.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Thu, 10 Apr 2008 10:48:29 +0800 (SGT)
Date: Thu, 10 Apr 2008 10:48:04 +0800
From: Zhijun Fu <Zhijun.Fu@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <18428.49039.503949.2406@gargle.gargle.HOWL>
Sender: Zhijun.Fu@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: Darren Reed <Darren.Reed@sun.com>, PSARC-EXT <psarc-ext@sun.com>
Reply-to: Zhijun.Fu@sun.com
Message-id: <47FD7FE4.8020601@Sun.COM>
MIME-version: 1.0
Content-type: multipart/alternative;
 boundary="Boundary_(ID_raMVzczwnOMJk1KJhR+HIQ)"
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
 <18428.49039.503949.2406@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.6 (X11/20071119)
Status: RO
Content-Length: 14412

This is a multi-part message in MIME format.

--Boundary_(ID_raMVzczwnOMJk1KJhR+HIQ)
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT

James Carlson wrote:
> Darren Reed writes:
>   
>> This case will extend PSARC/2005/334, by adding the ability to intercept
>> packets in MAC layer using the PFHooks infrastructure.
>>     
>
> Very minor nit: case title doesn't quite match the contents.  I was
> excited to see the name, because I'm working on MAC layer interception
> ... until I read that it was just a PFHooks extension for layer 2.
>
>   
>> Users can use ipf(1M) to add ethernet filtering rules in addition to IP 
>> filtering rules, the ethernet filtering rules are marked with "family ether".
>> Unlike IPv6, no special command line switch is required to load ethernet 
>> rules. And by default, ethernet rules should be put in /etc/ipf/ipf.conf.
>>     
>
> That seems strange.
>
> We currently have /etc/ipf/ipf.conf for IPv4 and the undocumented
> /etc/ipf/ipf6.conf for IPv6.  Why wouldn't we have /etc/ipf/ipfl2.conf
> (or some such) for L2-specific rules?
>
> Or if "family ether" is a good way to do this, why wouldn't we have
> "family inet" and "family inet6" and get rid of /etc/ipf/ipf6.conf?
>
> What's the intended direction?
>   
Darren has answered this so I'm taking the rest :-)
>> The layer 2 filtering functionality is disabled by default, to enable it,
>> add the following line to the top of ipf.conf:
>>
>> set intercept_layer2 true;
>>     
>
> If I have to say "family ether" in order to specify a rule that
> filters L2 packets, why do I need to give this extra command?  Doesn't
> the existence of at least one "family ether" rule mean that I intend
> to filter L2 packets as well (and thus I want interception turned on)?
>   
This behavior is consistent with ipfilter today.
Ipfilter will not be enabled automatically when you specify an ipfilter 
rule. Today the addition of ipfilter rules will fail if ipfilter is not 
enabled, which you need to enable by using "svcadm enable ipfilter" or 
"ipf -E". 
>   
>> To distinguish IP filter/NAT rules intended to be processed in layer 2
>> from the rest of ipfilter rules, an additional keyword "layer2" is added.
>> Those ipfilter rules to be processed in layer 2 are marked with "layer2",
>> so these rules won't be processed again when packets goes up to IP.
>>     
>
> If I have rules that have "layer2" set, then why do I need to specify
> "ip-head" or "ip-nat" in the "family ether" filter?  Shouldn't any
> rules with "layer2" set just _automatically_ match?
>   
> Or perhaps the question is this: why would I want to have rules
> specified as "layer2", but then specifically avoid sending some
> packets through those rules with "ip-head" or "ip-nat"?  
There are cases where users want to do IP Filtering / IP NAT in layer 2 
only for specified ethernet addresses, so we need to use the keyword 
"ip-head" and "ip-nat", to indicate that we do IP Filtering / IP NAT in 
layer 2 only for packets matching this "family ether" rule. And there 
are customer requests for this feature.
> If I did have
> such a case, why wouldn't I set up a "family ether" rule that
> specifies "quick" -- so that the rest of the "layer2"-tagged rules
> aren't examined at all?  That (using "quick" instead for the reverse
> sense) seems a lot clearer to me than "ip-head" or "ip-nat".
>   
I think "quick" means packets matching this rule will bypass the rules 
after it. But basically ethernet filtering rules and ipfilter rules are 
two sets of rules. So it may not be desirable to use "quick" for this.
Also, all ethernet filtering rules will be processed before all ipfilter 
rules, no matter which rule is specified earlier in the configuration 
file (or command line).

Say we have the following rules in ipf.conf, in the following order:

(1) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
    keep state group 10 layer2 
(2) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32
(3) rdr nge1 from 10.10.10.20/32 to 10.10.10.10/32 port = 7777 -> 10.10.10.10
    port 8888 tcp layer2
(4) pass in on nge1 family ether from 0:14:4f:8d:ae:23 to any ip-head 10 ip-nat


When a packet arrives at the layer 2 hook, rule (4) will be processed 
first, although it is specified last in the configuration file. In this 
case, using "quick" instead of "ip-head" and "ip-nat" will be 
mis-leading and non-intuitive, because the semantics of "quick" means 
you only bypass the rules *after* this rule.

And more, with "ip-head" you can direct packets to different ipfilter 
rule groups based on ethernet addresses, this provides more flexibility, 
so we can have:

pass in family ether from 0:14:4f:8d:ae:23 to any ip-head 10
pass in family ether from 0:14:4f:8d:ae:22 to any ip-head 20
pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
    keep state group 10 layer2 
pass in proto tcp from any to any group 20 layer2

> I suspect that most sane rule sets will start with something like
> this:
>
> 	pass in on nge1 family ether all ip-head ip-nat
>   
In some situation this might be true.
As mentioned above, in cases where users want to do IP Filtering / IP 
NAT in layer 2 only for specified ethernet addresses, ethernet filtering 
rules need to include the ethernet address matching part.
> ... so that the remaining "layer2" filter entries (who filters on
> explicit MAC addresses?
"family ether" rules filter on MAC addresses, and "layer2" rules 
filter/NAT on IP addresses.
> ) won't be confusing.
>
> (It seems to me that "family ether" and "layer2" are essentially the
> same thing; they're in lieu of putting these rules in a separate
> file.)
>   
Please see above. While putting these rules in a separate file seems 
doable, I failed to see how this can remove the need for these keywords. 
Please note that users can add rules on command line, and without 
"family ether" we have no way to distinguish ethernet filtering rules 
from ipfilter rules.

Say,
pass in family ether from any to any
pass in from any to any

------
Thanks,
Zhijun

-- 
#mdb -K
[0]> eri.prc.sun.com::walk staff s|::print staff_t s_email|
::grep .== Zhijun.Fu@Sun.COM|::eval <s=K|::print staff_t
Zhijun.Fu@Sun.COM, x84349
Network Virtualization & Performance Team,
Solaris Core Operating Systems
Since Jul 10,2006
[0]> :c


--Boundary_(ID_raMVzczwnOMJk1KJhR+HIQ)
Content-type: text/html; charset=ISO-8859-1
Content-transfer-encoding: 7BIT

<!DOCTYPE html PUBLIC "-//W3C//DTD HTML 4.01 Transitional//EN">
<html>
<head>
  <meta content="text/html;charset=ISO-8859-1" http-equiv="Content-Type">
</head>
<body bgcolor="#ffffff" text="#000000">
James Carlson wrote:
<blockquote cite="mid:18428.49039.503949.2406@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">Darren Reed writes:
  </pre>
  <blockquote type="cite">
    <pre wrap="">This case will extend PSARC/2005/334, by adding the ability to intercept
packets in MAC layer using the PFHooks infrastructure.
    </pre>
  </blockquote>
  <pre wrap=""><!---->
Very minor nit: case title doesn't quite match the contents.  I was
excited to see the name, because I'm working on MAC layer interception
... until I read that it was just a PFHooks extension for layer 2.

  </pre>
  <blockquote type="cite">
    <pre wrap="">Users can use ipf(1M) to add ethernet filtering rules in addition to IP 
filtering rules, the ethernet filtering rules are marked with "family ether".
Unlike IPv6, no special command line switch is required to load ethernet 
rules. And by default, ethernet rules should be put in /etc/ipf/ipf.conf.
    </pre>
  </blockquote>
  <pre wrap=""><!---->
That seems strange.

We currently have /etc/ipf/ipf.conf for IPv4 and the undocumented
/etc/ipf/ipf6.conf for IPv6.  Why wouldn't we have /etc/ipf/ipfl2.conf
(or some such) for L2-specific rules?

Or if "family ether" is a good way to do this, why wouldn't we have
"family inet" and "family inet6" and get rid of /etc/ipf/ipf6.conf?

What's the intended direction?
  </pre>
</blockquote>
Darren has answered this so I'm taking the rest<span
 class="moz-smiley-s1"><span> :-) </span></span>
<blockquote cite="mid:18428.49039.503949.2406@gargle.gargle.HOWL"
 type="cite">
  <blockquote type="cite">
    <pre wrap="">The layer 2 filtering functionality is disabled by default, to enable it,
add the following line to the top of ipf.conf:

set intercept_layer2 true;
    </pre>
  </blockquote>
  <pre wrap=""><!---->
If I have to say "family ether" in order to specify a rule that
filters L2 packets, why do I need to give this extra command?  Doesn't
the existence of at least one "family ether" rule mean that I intend
to filter L2 packets as well (and thus I want interception turned on)?
  </pre>
</blockquote>
This behavior is consistent with ipfilter today.<br>
Ipfilter will not be enabled automatically when you specify an ipfilter
rule. Today the addition of ipfilter rules will fail if ipfilter is not
enabled, which you need to enable by using "svcadm enable ipfilter" or
"ipf -E".&nbsp; <br>
<blockquote cite="mid:18428.49039.503949.2406@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">
  </pre>
  <blockquote type="cite">
    <pre wrap="">To distinguish IP filter/NAT rules intended to be processed in layer 2
from the rest of ipfilter rules, an additional keyword "layer2" is added.
Those ipfilter rules to be processed in layer 2 are marked with "layer2",
so these rules won't be processed again when packets goes up to IP.
    </pre>
  </blockquote>
  <pre wrap=""><!---->
If I have rules that have "layer2" set, then why do I need to specify
"ip-head" or "ip-nat" in the "family ether" filter?  Shouldn't any
rules with "layer2" set just _automatically_ match?
  </pre>
</blockquote>
<blockquote cite="mid:18428.49039.503949.2406@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">
Or perhaps the question is this: why would I want to have rules
specified as "layer2", but then specifically avoid sending some
packets through those rules with "ip-head" or "ip-nat"?  </pre>
</blockquote>
There are cases where users want to do IP Filtering / IP NAT in layer 2
only for specified ethernet addresses, so we need to use the keyword
"ip-head" and "ip-nat", to indicate that we do IP Filtering / IP NAT in
layer 2 only for packets matching this "family ether" rule. And there
are customer requests for this feature.<br>
<blockquote cite="mid:18428.49039.503949.2406@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">If I did have
such a case, why wouldn't I set up a "family ether" rule that
specifies "quick" -- so that the rest of the "layer2"-tagged rules
aren't examined at all?  That (using "quick" instead for the reverse
sense) seems a lot clearer to me than "ip-head" or "ip-nat".
  </pre>
</blockquote>
I think "quick" means packets matching this rule will bypass the rules
after it. But basically ethernet filtering rules and ipfilter rules are
two sets of rules. So it may not be desirable to use "quick" for this.<br>
Also, all ethernet filtering rules will be processed before all
ipfilter rules, no matter which rule is specified earlier in the
configuration file (or command line).<br>
<br>
Say we have the following rules in ipf.conf, in the following order:<br>
<pre wrap="">(1) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
    keep state group 10 layer2 
(2) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32
(3) rdr nge1 from 10.10.10.20/32 to 10.10.10.10/32 port = 7777 -&gt; 10.10.10.10
    port 8888 tcp layer2
(4) pass in on nge1 family ether from 0:14:4f:8d:ae:23 to any ip-head 10 ip-nat

</pre>
When a packet arrives at the layer 2 hook, rule (4) will be processed
first, although it is specified last in the configuration file. In this
case, using "quick" instead of "ip-head" and "ip-nat" will be
mis-leading and non-intuitive, because the semantics of "quick" means
you only bypass the rules <b>after</b> this rule.<br>
<br>
And more, with "ip-head" you can direct packets to different ipfilter
rule groups based on ethernet addresses, this provides more
flexibility, so we can have:<br>
<pre wrap="">pass in family ether from 0:14:4f:8d:ae:23 to any ip-head 10
pass in family ether from 0:14:4f:8d:ae:22 to any ip-head 20
pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
    keep state group 10 layer2 
pass in proto tcp from any to any group 20 layer2
</pre>
<blockquote cite="mid:18428.49039.503949.2406@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">
I suspect that most sane rule sets will start with something like
this:

	pass in on nge1 family ether all ip-head ip-nat
  </pre>
</blockquote>
In some situation this might be true. <br>
As mentioned above, in cases where users want to do IP Filtering / IP
NAT in layer 2
only for specified ethernet addresses, ethernet filtering rules need to
include the ethernet address matching part.<br>
<blockquote cite="mid:18428.49039.503949.2406@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">
... so that the remaining "layer2" filter entries (who filters on
explicit MAC addresses?</pre>
</blockquote>
"family ether" rules filter on MAC addresses, and "layer2" rules
filter/NAT on IP addresses.<br>
<blockquote cite="mid:18428.49039.503949.2406@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">) won't be confusing.

(It seems to me that "family ether" and "layer2" are essentially the
same thing; they're in lieu of putting these rules in a separate
file.)
  </pre>
</blockquote>
Please see above. While putting these rules in a separate file seems
doable, I failed to see how this can remove the need for these
keywords. Please note that users can add rules on command line, and
without "family ether" we have no way to distinguish ethernet filtering
rules from ipfilter rules.<br>
<br>
Say, <br>
pass in family ether from any to any<br>
pass in from any to any<br>
<br>
------<br>
Thanks,<br>
Zhijun<br>
<pre class="moz-signature" cols="72">-- 
#mdb -K
[0]&gt; eri.prc.sun.com::walk staff s|::print staff_t s_email|
::grep .== <a class="moz-txt-link-abbreviated" href="mailto:Zhijun.Fu@Sun.COM">Zhijun.Fu@Sun.COM</a>|::eval &lt;s=K|::print staff_t
<a class="moz-txt-link-abbreviated" href="mailto:Zhijun.Fu@Sun.COM">Zhijun.Fu@Sun.COM</a>, x84349
Network Virtualization &amp; Performance Team,
Solaris Core Operating Systems
Since Jul 10,2006
[0]&gt; :c
</pre>
</body>
</html>

--Boundary_(ID_raMVzczwnOMJk1KJhR+HIQ)--

From carlsonj@phorcys.east.sun.com Thu Apr 10 06:00:44 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m3AD0i2X001323
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 10 Apr 2008 06:00:44 -0700 (PDT)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m3AD0fsZ034705;
	Thu, 10 Apr 2008 07:00:42 -0600 (MDT)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ400I0T1H5KB00@nwk-avmta-2.sfbay.sun.com>; Thu,
 10 Apr 2008 06:00:41 -0700 (PDT)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ400H921H2YB80@nwk-avmta-2.sfbay.sun.com>; Thu,
 10 Apr 2008 06:00:39 -0700 (PDT)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m3AD0cXQ018312; Thu,
 10 Apr 2008 09:00:38 -0400 (EDT)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m3AD0c73018299; Thu,
 10 Apr 2008 09:00:38 -0400 (EDT)
Date: Thu, 10 Apr 2008 09:00:38 -0400
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <47FD7FE4.8020601@Sun.COM>
To: Zhijun.Fu@sun.com
Cc: Darren Reed <Darren.Reed@sun.com>, PSARC-EXT <psarc-ext@sun.com>
Message-id: <18430.3958.401815.539129@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
 <18428.49039.503949.2406@gargle.gargle.HOWL> <47FD7FE4.8020601@Sun.COM>
Status: RO
Content-Length: 7480

Zhijun Fu writes:
> James Carlson wrote:
> > If I have to say "family ether" in order to specify a rule that
> > filters L2 packets, why do I need to give this extra command?  Doesn't
> > the existence of at least one "family ether" rule mean that I intend
> > to filter L2 packets as well (and thus I want interception turned on)?
> >   
> This behavior is consistent with ipfilter today.
> Ipfilter will not be enabled automatically when you specify an ipfilter 
> rule. Today the addition of ipfilter rules will fail if ipfilter is not 
> enabled, which you need to enable by using "svcadm enable ipfilter" or 
> "ipf -E". 

No, it's *not* consistent.

All that I have to do today is to add rules into /etc/ipf/ipf.conf and
then enable filtering with "svcadm enable ipfilter".  I don't have to
do anything else.

This project proposes a new mechanism -- yet another hurdle to jump --
where I must add an "option" in the configuration file to tell the
system what the configuration file apparently already says.

Why do I need to specify "set intercept_layer2 true;"?

I think there's a misunderstanding here.  The reason that special "set
intercept_loopback true;" magic was added to the configuration file
was that old configuration files would be mis-interpreted by ipfilter
after an upgrade.  Configuration files that once didn't apply to local
traffic would suddenly start filtering that traffic unexpectedly when
the loopback filtering feature was added.

It was an incompatible change, so we wanted to have administrators
carefully consider how to update the rules before springing the change
on them.  That's why the flag is there.

I see *NO* such issue at all for the new L2 rules.  They don't change
the interpretation of any existing configuration files, because no
existing configuration file will have "family ether" or "layer2"
specified in it.

Thus, the new switch isn't necessary.

> > Or perhaps the question is this: why would I want to have rules
> > specified as "layer2", but then specifically avoid sending some
> > packets through those rules with "ip-head" or "ip-nat"?  
> There are cases where users want to do IP Filtering / IP NAT in layer 2 
> only for specified ethernet addresses, so we need to use the keyword 
> "ip-head" and "ip-nat", to indicate that we do IP Filtering / IP NAT in 
> layer 2 only for packets matching this "family ether" rule. And there 
> are customer requests for this feature.

There may be such cases, but that's exactly what the existing "group"
feature in IP filter is designed to handle.  It allows you to specify
that packets matching certain criteria are to be further processed by
a subset of the rules.

This new "ip-head" and "ip-nat" feature does the same thing, but less
flexibly.

> > If I did have
> > such a case, why wouldn't I set up a "family ether" rule that
> > specifies "quick" -- so that the rest of the "layer2"-tagged rules
> > aren't examined at all?  That (using "quick" instead for the reverse
> > sense) seems a lot clearer to me than "ip-head" or "ip-nat".
> >   
> I think "quick" means packets matching this rule will bypass the rules 
> after it.

That's correct.

> But basically ethernet filtering rules and ipfilter rules are 
> two sets of rules. So it may not be desirable to use "quick" for this.

I don't understand that concern.  Please elaborate on how introducing
these new keywords avoids some problem.  It looks to me like they do
essentially what the old keywords do -- just backwards.  The rule
matching is "opt in" rather than "opt out."

> Also, all ethernet filtering rules will be processed before all ipfilter 
> rules, no matter which rule is specified earlier in the configuration 
> file (or command line).

Sure; understood.

We're talking only about the way in which the "family ether" and
"layer2" rules work, which are all processed in the L2 data path.  I'm
not talking about any rules that lack those keywords.

That's not the issue.

> Say we have the following rules in ipf.conf, in the following order:
> 
> (1) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
>     keep state group 10 layer2 
> (2) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32
> (3) rdr nge1 from 10.10.10.20/32 to 10.10.10.10/32 port = 7777 -> 10.10.10.10
>     port 8888 tcp layer2
> (4) pass in on nge1 family ether from 0:14:4f:8d:ae:23 to any ip-head 10 ip-nat
> 
> 
> When a packet arrives at the layer 2 hook, rule (4) will be processed 
> first, although it is specified last in the configuration file. In this 
> case, using "quick" instead of "ip-head" and "ip-nat" will be 
> mis-leading and non-intuitive, because the semantics of "quick" means 
> you only bypass the rules *after* this rule.

To get that same behavior with the existing keywords (removing the
need for "ip-head" and "ip-nat"), I would use:

pass in on nge1 family ether from 0:14:4f:8d:ae:23 to any head 10
pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
   keep state group 10 layer2
pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32
rdr nge1 from 10.10.10.20/32 to 10.10.10.10/32 port = 7777 -> 10.10.10.10
  port 8888 tcp group 10 layer2

> And more, with "ip-head" you can direct packets to different ipfilter 
> rule groups based on ethernet addresses, this provides more flexibility, 
> so we can have:
> 
> pass in family ether from 0:14:4f:8d:ae:23 to any ip-head 10
> pass in family ether from 0:14:4f:8d:ae:22 to any ip-head 20
> pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
>     keep state group 10 layer2 
> pass in proto tcp from any to any group 20 layer2

I don't think that's necessary at all.  The rules would work fine
without "ip-head" (using just the existing "head" keyword) because of
the grouping rules.

> > I suspect that most sane rule sets will start with something like
> > this:
> >
> > 	pass in on nge1 family ether all ip-head ip-nat
> >   
> In some situation this might be true.
> As mentioned above, in cases where users want to do IP Filtering / IP 
> NAT in layer 2 only for specified ethernet addresses, ethernet filtering 
> rules need to include the ethernet address matching part.

I'm still missing why new keywords are needed.  The existing ones seem
to work just fine here.

> > (It seems to me that "family ether" and "layer2" are essentially the
> > same thing; they're in lieu of putting these rules in a separate
> > file.)
> >   
> Please see above. While putting these rules in a separate file seems 
> doable, I failed to see how this can remove the need for these keywords. 
> Please note that users can add rules on command line, and without 
> "family ether" we have no way to distinguish ethernet filtering rules 
> from ipfilter rules.

Yep; understood.  You'd end up with separate command line flags if you
did that, just as IPv6 does today.

My point in saying this is that the design looks incomplete.  It looks
like you've added L2 (Ethernet, really; nothing else seems to be
supported) to a single configuration file, but the equivalent work for
IPv6 hasn't been done yet, so the design is "mixed" and confusing as a
result.

All of the L2 examples you've provided use IPv4.  What would happen if
I needed to add L2+IPv6 rules?  Can I do that?

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Zhijun.Fu@sun.com Thu Apr 10 09:39:26 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m3AGdP0J008187
	for <psarc-ext@sac.sfbay.Sun.COM>; Thu, 10 Apr 2008 09:39:26 -0700 (PDT)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id m3AGdMQo006923
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Fri, 11 Apr 2008 00:39:24 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ400G0LBLN4B00@nwk-avmta-1.sfbay.Sun.COM> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Thu, 10 Apr 2008 09:39:23 -0700 (PDT)
Received: from sineb-mail-2.sun.com ([192.18.19.7])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ400EGNBLJOR30@nwk-avmta-1.sfbay.Sun.COM> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Thu,
 10 Apr 2008 09:39:20 -0700 (PDT)
Received: from fe-apac-06.sun.com
 (fe-apac-06.sun.com [192.18.19.177] (may be forged))
	by sineb-mail-2.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m3AGdnh0003300	for
 <psarc-ext@sun.com>; Thu, 10 Apr 2008 16:39:49 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JZ400D01BAQR500@mail-apac.sun.com> (original mail from Zhijun.Fu@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Fri,
 11 Apr 2008 00:38:46 +0800 (SGT)
Received: from [192.168.0.2] ([221.221.241.133])
 by mail-apac.sun.com (Sun Java System Messaging Server 6.2-6.01 (built Apr  3
 2006)) with ESMTPSA id <0JZ400BTHBKIRFLG@mail-apac.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Fri,
 11 Apr 2008 00:38:46 +0800 (SGT)
Date: Fri, 11 Apr 2008 00:40:30 +0800
From: Zhijun <Zhijun.Fu@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <18430.3958.401815.539129@gargle.gargle.HOWL>
Sender: Zhijun.Fu@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: Darren Reed <Darren.Reed@sun.com>, PSARC-EXT <psarc-ext@sun.com>
Message-id: <47FE42FE.3090308@Sun.COM>
MIME-version: 1.0
Content-type: multipart/alternative;
 boundary="Boundary_(ID_PbOCoKOdTtSIvMKO/mb7lg)"
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
 <18428.49039.503949.2406@gargle.gargle.HOWL> <47FD7FE4.8020601@Sun.COM>
 <18430.3958.401815.539129@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.12 (Windows/20080213)
Status: RO
Content-Length: 25855

This is a multi-part message in MIME format.

--Boundary_(ID_PbOCoKOdTtSIvMKO/mb7lg)
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT

James Carlson wrote:
> Zhijun Fu writes:
>   
>> James Carlson wrote:
>>     
>>> If I have to say "family ether" in order to specify a rule that
>>> filters L2 packets, why do I need to give this extra command?  Doesn't
>>> the existence of at least one "family ether" rule mean that I intend
>>> to filter L2 packets as well (and thus I want interception turned on)?
>>>   
>>>       
>> This behavior is consistent with ipfilter today.
>> Ipfilter will not be enabled automatically when you specify an ipfilter 
>> rule. Today the addition of ipfilter rules will fail if ipfilter is not 
>> enabled, which you need to enable by using "svcadm enable ipfilter" or 
>> "ipf -E". 
>>     
James,
Many thanks for helping review the case,  my input inline...
> No, it's *not* consistent.
>
> All that I have to do today is to add rules into /etc/ipf/ipf.conf and
> then enable filtering with "svcadm enable ipfilter".  I don't have to
> do anything else.
>
> This project proposes a new mechanism -- yet another hurdle to jump --
> where I must add an "option" in the configuration file to tell the
> system what the configuration file apparently already says.
>
> Why do I need to specify "set intercept_layer2 true;"?
>
> I think there's a misunderstanding here.  The reason that special "set
> intercept_loopback true;" magic was added to the configuration file
> was that old configuration files would be mis-interpreted by ipfilter
> after an upgrade.  Configuration files that once didn't apply to local
> traffic would suddenly start filtering that traffic unexpectedly when
> the loopback filtering feature was added.
>
> It was an incompatible change, so we wanted to have administrators
> carefully consider how to update the rules before springing the change
> on them.  That's why the flag is there.
>   
OK,  understood. Thank you for your explanation about this.
> I see *NO* such issue at all for the new L2 rules.  They don't change
> the interpretation of any existing configuration files, because no
> existing configuration file will have "family ether" or "layer2"
> specified in it.
>   
Agreed. L2 rules won't have the incompatible issues that loopback rules 
have.
> Thus, the new switch isn't necessary.
>   
I think the switch is still needed, for a different reason, so let me 
explain here:

Basically we have 3 options for this,
* have layer 2 filtering enabled by default, so we don't need to add a 
new switch to enable it
* disable layer 2 filtering by default, and enable it when users first 
adds a rule contains keyword "family ether"
* disable layer 2 filtering by default, and specify "set 
intercept_layer2 true;" in /etc/ipf/ipf.conf to enable it, and this is 
the one we proposed

With the 1st option, having layer 2 filtering enabled by default will 
have performance impact, even if there're no layer 2 rules configured, 
because additional processing is needed when layer 2 filtering is 
enabled, so this option cannot be used.
With the 2nd option, we'll enable the layer 2 filtering when the first 
"family ether" rule is added, and disable layer 2 filtering when the 
last "family ether" rule is removed, this has the following issues
- We need to add check for each addition/removal of "family ether" 
rules, for each addition/removal, we need to check whether this is the 
first rule added, or the last rule removed, I'm not sure whether this is 
desirable
- Enabling layer 2 filtering involves registering hooks, today for 
ipfilter, both IPv4 and IPv6 hooks are registered when ipfilter is 
enabled, and unregistered when ipfilter is disabled; for loopback 
filtering, the case is the same. So I think for layer 2 filtering, 
registering/unregistering hooks when a rule is added/removed instead of 
the feature (layer 2 filtering) enabled/disabled sounds inconsistent 
with the existing behavior
- Today, with ipfilter you need to enable the ipfilter service before 
you can add rules. If we follow the logic of option 2, why not we change 
ipfilter and let it get enabled automatically when the first ipfilter 
rule is added? But this is not what ipfilter behaves in neveda.

I understand optioin 3 may not be a perfect solution in all aspects, but 
if we need a trade-off, sounds like it is a better choice than the rest.
>>> Or perhaps the question is this: why would I want to have rules
>>> specified as "layer2", but then specifically avoid sending some
>>> packets through those rules with "ip-head" or "ip-nat"?  
>>>       
>> There are cases where users want to do IP Filtering / IP NAT in layer 2 
>> only for specified ethernet addresses, so we need to use the keyword 
>> "ip-head" and "ip-nat", to indicate that we do IP Filtering / IP NAT in 
>> layer 2 only for packets matching this "family ether" rule. And there 
>> are customer requests for this feature.
>>     
>
> There may be such cases, but that's exactly what the existing "group"
> feature in IP filter is designed to handle.  It allows you to specify
> that packets matching certain criteria are to be further processed by
> a subset of the rules.
>   
Right. And we're using "group" together with "ip-head".
> This new "ip-head" and "ip-nat" feature does the same thing, 
"ip-head" does the similar thing for "head", and "ip-nat" is different.
Please note the group feature is only available for IP Filtering rules 
today,  there're no group support for IP NAT rules. So "ip-nat" is 
irrelevant with the "group" feature.
> but less
> flexibly.
>   
>>> If I did have
>>> such a case, why wouldn't I set up a "family ether" rule that
>>> specifies "quick" -- so that the rest of the "layer2"-tagged rules
>>> aren't examined at all?  That (using "quick" instead for the reverse
>>> sense) seems a lot clearer to me than "ip-head" or "ip-nat".
>>>   
>>>       
>> I think "quick" means packets matching this rule will bypass the rules 
>> after it.
>>     
>
> That's correct.
>
>   
>> But basically ethernet filtering rules and ipfilter rules are 
>> two sets of rules. So it may not be desirable to use "quick" for this.
>>     
>
> I don't understand that concern.  Please elaborate on how introducing
> these new keywords avoids some problem.  It looks to me like they do
> essentially what the old keywords do -- just backwards.  The rule
> matching is "opt in" rather than "opt out."
>   
I meant use "quick" here may not be desirable because the issues 
mentioned below. Sorry for the confusion.
>   
>> Also, all ethernet filtering rules will be processed before all ipfilter 
>> rules, no matter which rule is specified earlier in the configuration 
>> file (or command line).
>>     
>
> Sure; understood.
>
> We're talking only about the way in which the "family ether" and
> "layer2" rules work, which are all processed in the L2 data path.  
> I'm
> not talking about any rules that lack those keywords.
>   
> That's not the issue.
>   
I was talking about using "quick" instead of "ip-head" "ip-nat" is not 
desirable, because the order rules listed in the configuration file may 
not be the same order that rules get processed, so using "quick" can be 
misleading.
>> Say we have the following rules in ipf.conf, in the following order:
>>
>> (1) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
>>     keep state group 10 layer2 
>> (2) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32
>> (3) rdr nge1 from 10.10.10.20/32 to 10.10.10.10/32 port = 7777 -> 10.10.10.10
>>     port 8888 tcp layer2
>> (4) pass in on nge1 family ether from 0:14:4f:8d:ae:23 to any ip-head 10 ip-nat
>>
>>
>> When a packet arrives at the layer 2 hook, rule (4) will be processed 
>> first, although it is specified last in the configuration file. In this 
>> case, using "quick" instead of "ip-head" and "ip-nat" will be 
>> mis-leading and non-intuitive, because the semantics of "quick" means 
>> you only bypass the rules *after* this rule.
>>     
>
> To get that same behavior with the existing keywords (removing the
> need for "ip-head" and "ip-nat"), I would use:
>
> pass in on nge1 family ether from 0:14:4f:8d:ae:23 to any head 10
> pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
>    keep state group 10 layer2
> pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32
> rdr nge1 from 10.10.10.20/32 to 10.10.10.10/32 port = 7777 -> 10.10.10.10
>   port 8888 tcp group 10 layer2
>   
Today we don't have group support for IP NAT rules in neveda.
Adding group support for IP NAT rules can be a way to solve this, and 
using "ip-nat" is another way, and requires less changes.
>> And more, with "ip-head" you can direct packets to different ipfilter 
>> rule groups based on ethernet addresses, this provides more flexibility, 
>> so we can have:
>>
>> pass in family ether from 0:14:4f:8d:ae:23 to any ip-head 10
>> pass in family ether from 0:14:4f:8d:ae:22 to any ip-head 20
>> pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
>>     keep state group 10 layer2 
>> pass in proto tcp from any to any group 20 layer2
>>     
>
> I don't think that's necessary at all.  The rules would work fine
> without "ip-head" (using just the existing "head" keyword) because of
> the grouping rules.
>   
I may need to think about this more.
>>> I suspect that most sane rule sets will start with something like
>>> this:
>>>
>>> 	pass in on nge1 family ether all ip-head ip-nat
>>>   
>>>       
>> In some situation this might be true.
>> As mentioned above, in cases where users want to do IP Filtering / IP 
>> NAT in layer 2 only for specified ethernet addresses, ethernet filtering 
>> rules need to include the ethernet address matching part.
>>     
>
> I'm still missing why new keywords are needed.  The existing ones seem
> to work just fine here.
>
>   
>>> (It seems to me that "family ether" and "layer2" are essentially the
>>> same thing; they're in lieu of putting these rules in a separate
>>> file.)
>>>   
>>>       
>> Please see above. While putting these rules in a separate file seems 
>> doable, I failed to see how this can remove the need for these keywords. 
>> Please note that users can add rules on command line, and without 
>> "family ether" we have no way to distinguish ethernet filtering rules 
>> from ipfilter rules.
>>     
>
> Yep; understood.  You'd end up with separate command line flags if you
> did that, just as IPv6 does today.
>
> My point in saying this is that the design looks incomplete.  It looks
> like you've added L2 (Ethernet, really; nothing else seems to be
> supported) to a single configuration file, but the equivalent work for
> IPv6 hasn't been done yet, so the design is "mixed" and confusing as a
> result.
>   
As Darren has already explained, /etc/ipf/ipf6.conf was introduced as an 
obsolete interface for a historic reason. I fully agree that it is good 
and desirable to have all IPv4, IPv6 and ether rules consistent (maybe 
put altogether in ipf.conf), I'm just not sure if it is needed to be 
done in this project. As this project is only focused on layer 2 
filtering and doesn't touch anything about IPv6. So probably this can be 
addressed via a separate bug/RFE?
> All of the L2 examples you've provided use IPv4.  What would happen if
> I needed to add L2+IPv6 rules?  Can I do that?
>   
Sure. Just add the "layer2" keyword to the IPv6 rules, the usage is the 
same as IPv4 rules. :-)

Thanks again,
Zhijun


--Boundary_(ID_PbOCoKOdTtSIvMKO/mb7lg)
Content-type: text/html; charset=ISO-8859-1
Content-transfer-encoding: 7BIT

<!DOCTYPE html PUBLIC "-//W3C//DTD HTML 4.01 Transitional//EN">
<html>
<head>
  <meta content="text/html;charset=ISO-8859-1" http-equiv="Content-Type">
  <title></title>
</head>
<body bgcolor="#ffffff" text="#000000">
James Carlson wrote:
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">Zhijun Fu writes:
  </pre>
  <blockquote type="cite">
    <pre wrap="">James Carlson wrote:
    </pre>
    <blockquote type="cite">
      <pre wrap="">If I have to say "family ether" in order to specify a rule that
filters L2 packets, why do I need to give this extra command?  Doesn't
the existence of at least one "family ether" rule mean that I intend
to filter L2 packets as well (and thus I want interception turned on)?
  
      </pre>
    </blockquote>
    <pre wrap="">This behavior is consistent with ipfilter today.
Ipfilter will not be enabled automatically when you specify an ipfilter 
rule. Today the addition of ipfilter rules will fail if ipfilter is not 
enabled, which you need to enable by using "svcadm enable ipfilter" or 
"ipf -E". 
    </pre>
  </blockquote>
</blockquote>
James, <br>
Many thanks for helping review the case,&nbsp; my input inline...<br>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">No, it's *not* consistent.

All that I have to do today is to add rules into /etc/ipf/ipf.conf and
then enable filtering with "svcadm enable ipfilter".  I don't have to
do anything else.

This project proposes a new mechanism -- yet another hurdle to jump --
where I must add an "option" in the configuration file to tell the
system what the configuration file apparently already says.

Why do I need to specify "set intercept_layer2 true;"?

I think there's a misunderstanding here.  The reason that special "set
intercept_loopback true;" magic was added to the configuration file
was that old configuration files would be mis-interpreted by ipfilter
after an upgrade.  Configuration files that once didn't apply to local
traffic would suddenly start filtering that traffic unexpectedly when
the loopback filtering feature was added.

It was an incompatible change, so we wanted to have administrators
carefully consider how to update the rules before springing the change
on them.  That's why the flag is there.
  </pre>
</blockquote>
OK,&nbsp; understood. Thank you for your explanation about this.<br>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">
I see *NO* such issue at all for the new L2 rules.  They don't change
the interpretation of any existing configuration files, because no
existing configuration file will have "family ether" or "layer2"
specified in it.
  </pre>
</blockquote>
Agreed. L2 rules won't have the incompatible issues that loopback rules
have.<br>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">
Thus, the new switch isn't necessary.
  </pre>
</blockquote>
I think the switch is still needed, for a different reason, so let me
explain here:<br>
<br>
Basically we have 3 options for this,<br>
* have layer 2 filtering enabled by default, so we don't need to add a
new switch to enable it<br>
* disable layer 2 filtering by default, and enable it when users first
adds a rule contains keyword "family ether"<br>
* disable layer 2 filtering by default, and specify "set
intercept_layer2 true;" in /etc/ipf/ipf.conf to enable it, and this is
the one we proposed<br>
<br>
With the 1st option, having layer 2 filtering enabled by default will
have performance impact, even if there're no layer 2 rules configured,
because additional processing is needed when layer 2 filtering is
enabled, so this option cannot be used.<br>
With the 2nd option, we'll enable the layer 2 filtering when the first
"family ether" rule is added, and disable layer 2 filtering when the
last "family ether" rule is removed, this has the following issues<br>
- We need to add check for each addition/removal of "family ether"
rules, for each addition/removal, we need to check whether this is the
first rule added, or the last rule removed, I'm not sure whether this
is desirable<br>
- Enabling layer 2 filtering involves registering hooks, today for
ipfilter, both IPv4 and IPv6 hooks are registered when ipfilter is
enabled, and unregistered when ipfilter is disabled; for loopback
filtering, the case is the same. So I think for layer 2 filtering,
registering/unregistering hooks when a rule is added/removed instead of
the feature (layer 2 filtering) enabled/disabled sounds inconsistent
with the existing behavior<br>
- Today, with ipfilter you need to enable the ipfilter service before
you can add rules. If we follow the logic of option 2, why not we
change ipfilter and let it get enabled automatically when the first
ipfilter rule is added? But this is not what ipfilter behaves in neveda.<br>
<br>
I understand optioin 3 may not be a perfect solution in all aspects,
but if we need a trade-off, sounds like it is a better choice than the
rest.<br>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <blockquote type="cite">
    <blockquote type="cite">
      <pre wrap="">Or perhaps the question is this: why would I want to have rules
specified as "layer2", but then specifically avoid sending some
packets through those rules with "ip-head" or "ip-nat"?  
      </pre>
    </blockquote>
    <pre wrap="">There are cases where users want to do IP Filtering / IP NAT in layer 2 
only for specified ethernet addresses, so we need to use the keyword 
"ip-head" and "ip-nat", to indicate that we do IP Filtering / IP NAT in 
layer 2 only for packets matching this "family ether" rule. And there 
are customer requests for this feature.
    </pre>
  </blockquote>
  <pre wrap=""><!---->
There may be such cases, but that's exactly what the existing "group"
feature in IP filter is designed to handle.  It allows you to specify
that packets matching certain criteria are to be further processed by
a subset of the rules.
  </pre>
</blockquote>
Right. And we're using "group" together with "ip-head".<br>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">
This new "ip-head" and "ip-nat" feature does the same thing, </pre>
</blockquote>
"ip-head" does the similar thing for "head", and "ip-nat" is different.<br>
Please note the group feature is only available for IP Filtering rules
today,&nbsp; there're no group support for IP NAT rules. So "ip-nat" is
irrelevant with the "group" feature.<br>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">but less
flexibly.
  </pre>
</blockquote>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <blockquote type="cite">
    <blockquote type="cite">
      <pre wrap="">If I did have
such a case, why wouldn't I set up a "family ether" rule that
specifies "quick" -- so that the rest of the "layer2"-tagged rules
aren't examined at all?  That (using "quick" instead for the reverse
sense) seems a lot clearer to me than "ip-head" or "ip-nat".
  
      </pre>
    </blockquote>
    <pre wrap="">I think "quick" means packets matching this rule will bypass the rules 
after it.
    </pre>
  </blockquote>
  <pre wrap=""><!---->
That's correct.

  </pre>
  <blockquote type="cite">
    <pre wrap="">But basically ethernet filtering rules and ipfilter rules are 
two sets of rules. So it may not be desirable to use "quick" for this.
    </pre>
  </blockquote>
  <pre wrap=""><!---->
I don't understand that concern.  Please elaborate on how introducing
these new keywords avoids some problem.  It looks to me like they do
essentially what the old keywords do -- just backwards.  The rule
matching is "opt in" rather than "opt out."
  </pre>
</blockquote>
I meant use "quick" here may not be desirable because the issues
mentioned below. Sorry for the confusion.<br>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">
  </pre>
  <blockquote type="cite">
    <pre wrap="">Also, all ethernet filtering rules will be processed before all ipfilter 
rules, no matter which rule is specified earlier in the configuration 
file (or command line).
    </pre>
  </blockquote>
  <pre wrap=""><!---->
Sure; understood.

We're talking only about the way in which the "family ether" and
"layer2" rules work, which are all processed in the L2 data path.  </pre>
</blockquote>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">I'm
not talking about any rules that lack those keywords.
  </pre>
</blockquote>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">
That's not the issue.
  </pre>
</blockquote>
I was talking about using "quick" instead of "ip-head" "ip-nat" is not
desirable, because the order rules listed in the configuration file may
not be the same order that rules get processed, so using "quick" can be
misleading.<br>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <blockquote type="cite">
    <pre wrap="">Say we have the following rules in ipf.conf, in the following order:

(1) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
    keep state group 10 layer2 
(2) pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32
(3) rdr nge1 from 10.10.10.20/32 to 10.10.10.10/32 port = 7777 -&gt; 10.10.10.10
    port 8888 tcp layer2
(4) pass in on nge1 family ether from 0:14:4f:8d:ae:23 to any ip-head 10 ip-nat


When a packet arrives at the layer 2 hook, rule (4) will be processed 
first, although it is specified last in the configuration file. In this 
case, using "quick" instead of "ip-head" and "ip-nat" will be 
mis-leading and non-intuitive, because the semantics of "quick" means 
you only bypass the rules *after* this rule.
    </pre>
  </blockquote>
  <pre wrap=""><!---->
To get that same behavior with the existing keywords (removing the
need for "ip-head" and "ip-nat"), I would use:

pass in on nge1 family ether from 0:14:4f:8d:ae:23 to any head 10
pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
   keep state group 10 layer2
pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32
rdr nge1 from 10.10.10.20/32 to 10.10.10.10/32 port = 7777 -&gt; 10.10.10.10
  port 8888 tcp group 10 layer2
  </pre>
</blockquote>
Today we don't have group support for IP NAT rules in neveda. <br>
Adding group support for IP NAT rules can be a way to solve this, and
using "ip-nat" is another way, and requires less changes.<br>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <blockquote type="cite">
    <pre wrap="">And more, with "ip-head" you can direct packets to different ipfilter 
rule groups based on ethernet addresses, this provides more flexibility, 
so we can have:

pass in family ether from 0:14:4f:8d:ae:23 to any ip-head 10
pass in family ether from 0:14:4f:8d:ae:22 to any ip-head 20
pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
    keep state group 10 layer2 
pass in proto tcp from any to any group 20 layer2
    </pre>
  </blockquote>
  <pre wrap=""><!---->
I don't think that's necessary at all.  The rules would work fine
without "ip-head" (using just the existing "head" keyword) because of
the grouping rules.
  </pre>
</blockquote>
I may need to think about this more.<br>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <blockquote type="cite">
    <blockquote type="cite">
      <pre wrap="">I suspect that most sane rule sets will start with something like
this:

	pass in on nge1 family ether all ip-head ip-nat
  
      </pre>
    </blockquote>
    <pre wrap="">In some situation this might be true.
As mentioned above, in cases where users want to do IP Filtering / IP 
NAT in layer 2 only for specified ethernet addresses, ethernet filtering 
rules need to include the ethernet address matching part.
    </pre>
  </blockquote>
  <pre wrap=""><!---->
I'm still missing why new keywords are needed.  The existing ones seem
to work just fine here.

  </pre>
  <blockquote type="cite">
    <blockquote type="cite">
      <pre wrap="">(It seems to me that "family ether" and "layer2" are essentially the
same thing; they're in lieu of putting these rules in a separate
file.)
  
      </pre>
    </blockquote>
    <pre wrap="">Please see above. While putting these rules in a separate file seems 
doable, I failed to see how this can remove the need for these keywords. 
Please note that users can add rules on command line, and without 
"family ether" we have no way to distinguish ethernet filtering rules 
from ipfilter rules.
    </pre>
  </blockquote>
  <pre wrap=""><!---->
Yep; understood.  You'd end up with separate command line flags if you
did that, just as IPv6 does today.

My point in saying this is that the design looks incomplete.  It looks
like you've added L2 (Ethernet, really; nothing else seems to be
supported) to a single configuration file, but the equivalent work for
IPv6 hasn't been done yet, so the design is "mixed" and confusing as a
result.
  </pre>
</blockquote>
As Darren has already explained, /etc/ipf/ipf6.conf was introduced as
an obsolete interface for a historic reason. I fully agree that it is
good and desirable to have all IPv4, IPv6 and ether rules consistent
(maybe put altogether in ipf.conf), I'm just not sure if it is needed
to be done in this project. As this project is only focused on layer 2
filtering and doesn't touch anything about IPv6. So probably this can
be addressed via a separate bug/RFE?<br>
<blockquote cite="mid:18430.3958.401815.539129@gargle.gargle.HOWL"
 type="cite">
  <pre wrap="">
All of the L2 examples you've provided use IPv4.  What would happen if
I needed to add L2+IPv6 rules?  Can I do that?
  </pre>
</blockquote>
Sure. Just add the "layer2" keyword to the IPv6 rules, the usage is the
same as IPv4 rules.<span class="moz-smiley-s1"><span> :-) </span></span><br>
<br>
Thanks again,<br>
Zhijun<br>
<br>
</body>
</html>

--Boundary_(ID_PbOCoKOdTtSIvMKO/mb7lg)--

From Darren.Moffat@Sun.COM Thu Apr 10 09:46:23 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m3AGkM7r008222
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 10 Apr 2008 09:46:23 -0700 (PDT)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id m3AGkKmS023520
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Thu, 10 Apr 2008 17:46:21 +0100 (BST)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ400207BX9IJ00@nwk-avmta-2.sfbay.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Thu, 10 Apr 2008 09:46:21 -0700 (PDT)
Received: from gmp-eb-inf-1.sun.com ([192.18.6.21])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ400KAFBX77IC0@nwk-avmta-2.sfbay.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Thu,
 10 Apr 2008 09:46:19 -0700 (PDT)
Received: from fe-emea-09.sun.com (gmp-eb-lb-2-fe1.eu.sun.com [192.18.6.10])
	by gmp-eb-inf-1.sun.com (8.13.7+Sun/8.12.9) with ESMTP id m3AGkI0v007375	for
 <psarc-ext@sun.com>; Thu, 10 Apr 2008 16:46:18 +0000 (GMT)
Received: from conversion-daemon.fe-emea-09.sun.com by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0JZ400L01B30C300@fe-emea-09.sun.com>
 (original mail from Darren.Moffat@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Thu,
 10 Apr 2008 17:46:18 +0100 (BST)
Received: from [129.156.173.199] by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0JZ400EEEBWN0N10@fe-emea-09.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Thu, 10 Apr 2008 17:46:00 +0100 (BST)
Date: Thu, 10 Apr 2008 17:45:59 +0100
From: Darren J Moffat <Darren.Moffat@Sun.COM>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <47FE42FE.3090308@Sun.COM>
Sender: Darren.Moffat@Sun.COM
To: Zhijun <Zhijun.Fu@Sun.COM>
Cc: James Carlson <James.D.Carlson@Sun.COM>, PSARC-EXT <psarc-ext@Sun.COM>,
        Darren Reed <Darren.Reed@Sun.COM>
Message-id: <47FE4447.10400@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
 <18428.49039.503949.2406@gargle.gargle.HOWL> <47FD7FE4.8020601@Sun.COM>
 <18430.3958.401815.539129@gargle.gargle.HOWL> <47FE42FE.3090308@Sun.COM>
User-Agent: Thunderbird 2.0.0.9 (X11/20080225)
Status: RO
Content-Length: 887

There has been mention of a possible performance impact several times 
but no evidence presented to prove that.  Is there data that shows there 
*is* performance impact or is there just an assumption that there would be ?

If there is evidence please provide it, if there isn't then please do 
the testing.

I think this case should be considered to be in "waiting need spec" 
until we know for sure if there actually is a performance impact or not.

As for why ipfilter is disabled by default that dates back to when the 
pfil module was used and that had a much more significant impact than 
anything the hooks could be doing because it impacted Fireengine.   I 
think (though not this case) it would be worth re considering if 
ipfilter should always be enabled (and again not this case) if a default 
rule set should be provided (SBD actually wanted to do this).

--
Darren J Moffat

From carlsonj@phorcys.east.sun.com Thu Apr 10 10:15:56 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m3AHFuL8008998
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 10 Apr 2008 10:15:56 -0700 (PDT)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m3AHFsJ6058040;
	Thu, 10 Apr 2008 11:15:54 -0600 (MDT)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ40030DDAICN00@nwk-avmta-2.sfbay.sun.com>; Thu,
 10 Apr 2008 10:15:54 -0700 (PDT)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ400KQ6DAG7VE0@nwk-avmta-2.sfbay.sun.com>; Thu,
 10 Apr 2008 10:15:53 -0700 (PDT)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2) with ESMTP id m3AHFq4F027516; Thu,
 10 Apr 2008 13:15:52 -0400 (EDT)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.2+Sun/8.14.2/Submit) id m3AHFq10027513; Thu,
 10 Apr 2008 13:15:52 -0400 (EDT)
Date: Thu, 10 Apr 2008 13:15:52 -0400
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <47FE42FE.3090308@Sun.COM>
To: Zhijun <Zhijun.Fu@sun.com>
Cc: Darren Reed <Darren.Reed@sun.com>, PSARC-EXT <psarc-ext@sun.com>
Message-id: <18430.19272.702199.86531@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
 <18428.49039.503949.2406@gargle.gargle.HOWL> <47FD7FE4.8020601@Sun.COM>
 <18430.3958.401815.539129@gargle.gargle.HOWL> <47FE42FE.3090308@Sun.COM>
Status: RO
Content-Length: 9480

Zhijun writes:
> Basically we have 3 options for this,
> * have layer 2 filtering enabled by default, so we don't need to add a 
> new switch to enable it
> * disable layer 2 filtering by default, and enable it when users first 
> adds a rule contains keyword "family ether"
> * disable layer 2 filtering by default, and specify "set 
> intercept_layer2 true;" in /etc/ipf/ipf.conf to enable it, and this is 
> the one we proposed
> 
> With the 1st option, having layer 2 filtering enabled by default will 
> have performance impact, even if there're no layer 2 rules configured, 
> because additional processing is needed when layer 2 filtering is 
> enabled, so this option cannot be used.

The performance impact is present only if the 'ipfilter' service is
enabled, which it's not.  And when you enable that service, you
already get an effect on IP, so having "more" shouldn't necessarily be
a problem: those using IP filter have already agreed to pay a
performance price for doing so.

Do you have any numbers to suggest that those who've intentionally
enabled IP filter would be upset at paying a slightly default higher
price in order to support L2 rules as well?  I think that's your
assertion, but I don't see why it'd be true.

However, that's not the option I'm arguing for anyway.

> With the 2nd option, we'll enable the layer 2 filtering when the first 
> "family ether" rule is added, and disable layer 2 filtering when the 
> last "family ether" rule is removed, this has the following issues
> - We need to add check for each addition/removal of "family ether" 
> rules, for each addition/removal, we need to check whether this is the 
> first rule added, or the last rule removed, I'm not sure whether this is 
> desirable

Actually, there are two variants that are possible:

	- Keep a count -- a single simple integer -- of all of the
          family ether and layer2 rules that are present.  If that
          count is not zero, then enable.  If it's zero, then disable.

	- Have a single *internal* flag.  If any "family ether" or
          "layer2" rule is inserted, then the flag is set.  The flag
          is reset only when all of the rules are flushed.  Users can
          then 'recover' from the overhead problem by doing "svcadm
          restart ipfilter".

Both of these appear to be simple and don't involve any extra
administrator-controlled flags.  That second one is *equivalent* to
what you're proposing, except that it doesn't require extra work by
the administrator.

> - Enabling layer 2 filtering involves registering hooks, today for 
> ipfilter, both IPv4 and IPv6 hooks are registered when ipfilter is 
> enabled, and unregistered when ipfilter is disabled; for loopback 
> filtering, the case is the same. So I think for layer 2 filtering, 
> registering/unregistering hooks when a rule is added/removed instead of 
> the feature (layer 2 filtering) enabled/disabled sounds inconsistent 
> with the existing behavior

Think ahead to when both IPv4 and IPv6 are in the same file, since
Darren has said that this is the intended goal.

It would be perfectly reasonable for us to register hooks for IPv4
only when IPv4 rules are present, and for IPv6 only when IPv6 is
present, so that overhead for one doesn't affect overhead for the
other.

In fact, taking it further still, it'd be perfectly reasonable to
refactor the hooks so that they're per-interface (rather than global),
and have the ipfilter service register hooks only on the interfaces
that are filtered -- so that only the filtered interfaces are saddled
with the overhead, and the other interfaces in the system are not.

That'd be a great improvement over the situation today, as inserting
filters for your 10Mbps management port slows down your 10Gbps data
port as well.

In any event, all of this is a purely internal design matter.  It's
not something that ought to have artifacts in the documented
configuration file.

Please don't push system architecture off onto the administrator.

> - Today, with ipfilter you need to enable the ipfilter service before 
> you can add rules. If we follow the logic of option 2, why not we change 
> ipfilter and let it get enabled automatically when the first ipfilter 
> rule is added? But this is not what ipfilter behaves in neveda.

That's not what I'm suggesting.

I'm suggesting that when the administrator puts rules into the file,
and then he enables the service, it should just work.  He should not
have to do anything else.  There should be no other switches or knobs
to tweak to make it "really enabled."

> I understand optioin 3 may not be a perfect solution in all aspects, but 
> if we need a trade-off, sounds like it is a better choice than the rest.

I think it's far worse.  It pushes off onto the administrator
something that the system ought to be handling itself.

I'm not talking about "svcadm enable ipfilter".  I still expect that
to be there.  I'm not asking for the ipfilter service to spring to
life magically when filters are added to the file.  Instead, I'm
asking that when it's enabled via SMF and when the rules are present,
those rules *WORK* and don't need extra tweaking.

Doing anything else is, to me, a regression in functionality.

> > There may be such cases, but that's exactly what the existing "group"
> > feature in IP filter is designed to handle.  It allows you to specify
> > that packets matching certain criteria are to be further processed by
> > a subset of the rules.
> >   
> Right. And we're using "group" together with "ip-head".
> > This new "ip-head" and "ip-nat" feature does the same thing, 
> "ip-head" does the similar thing for "head", and "ip-nat" is different.
> Please note the group feature is only available for IP Filtering rules 
> today,  there're no group support for IP NAT rules. So "ip-nat" is 
> irrelevant with the "group" feature.

That's perhaps a bug, but irrelevant.

You still haven't answered what "ip-head" does that "head" doesn't
already do.

> I was talking about using "quick" instead of "ip-head" "ip-nat" is not 
> desirable, because the order rules listed in the configuration file may 
> not be the same order that rules get processed, so using "quick" can be 
> misleading.

The rules in the file are always processed in the order that they're
inserted into the lists, as documented in the ipf.conf man page.

Order has always been significant.  I think you're confusing the
"group" feature and the "head" keyword with other bits -- it'd be
confusing indeed if L2 processing were forced to be always at the end
of the file or if it didn't follow the normal order-sensitive behavior
that IP Filter has always had.

> > pass in on nge1 family ether from 0:14:4f:8d:ae:23 to any head 10
> > pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32 port = 8888
> >    keep state group 10 layer2
> > pass in proto tcp from 10.10.10.20/32 to 10.10.10.10/32
> > rdr nge1 from 10.10.10.20/32 to 10.10.10.10/32 port = 7777 -> 10.10.10.10
> >   port 8888 tcp group 10 layer2
> >   
> Today we don't have group support for IP NAT rules in neveda.
> Adding group support for IP NAT rules can be a way to solve this, and 
> using "ip-nat" is another way, and requires less changes.

I suspect you can *still* do this without resorting to "ip-nat".

pass in on nge1 family ether from ! 0:14:4f:8d:ae:23 to any quick
rdr nge1 from 10.10.10.20/32 to 10.10.10.10/32 port = 7777 -> 10.10.10.10
  port 8888 tcp layer2

(Mixing ipf.conf and ipnat.conf syntax like this is pretty confusing
... is that part of this case ... ?)

> > My point in saying this is that the design looks incomplete.  It looks
> > like you've added L2 (Ethernet, really; nothing else seems to be
> > supported) to a single configuration file, but the equivalent work for
> > IPv6 hasn't been done yet, so the design is "mixed" and confusing as a
> > result.
> >   
> As Darren has already explained, /etc/ipf/ipf6.conf was introduced as an 
> obsolete interface for a historic reason. I fully agree that it is good 
> and desirable to have all IPv4, IPv6 and ether rules consistent (maybe 
> put altogether in ipf.conf), I'm just not sure if it is needed to be 
> done in this project. As this project is only focused on layer 2 
> filtering and doesn't touch anything about IPv6. So probably this can be 
> addressed via a separate bug/RFE?

Possibly.  If we can agree on what the architecture should look like.
I don't think we're at that point yet.

> > All of the L2 examples you've provided use IPv4.  What would happen if
> > I needed to add L2+IPv6 rules?  Can I do that?
> >   
> Sure. Just add the "layer2" keyword to the IPv6 rules, the usage is the 
> same as IPv4 rules. :-)

I don't think I follow.

The IPv6 rules are currently in a different file, because they need to
be processed with a special ipf command line flag.

So, is a "family ether" rule legal in ipf6.conf?  Since all of the L2
rules are undistinguished with respect to IPv4 and IPv6, what actual
order of L2 rules (in the kernel) do I get if I put these into both
ipf.conf and ipf6.conf?  (Again, I care about order because it's
significant for matching rules.)

These things seem to be in conflict, which is why I was pointing out
(above) that it doesn't appear to be complete.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Darren.Reed@sun.com Thu Apr 10 14:40:56 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m3ALeu5C017770
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 10 Apr 2008 14:40:56 -0700 (PDT)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id m3ALeqYh000921
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Thu, 10 Apr 2008 15:40:56 -0600 (MDT)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ400L07PK7HA00@nwk-avmta-1.sfbay.Sun.COM> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Thu, 10 Apr 2008 14:40:55 -0700 (PDT)
Received: from sineb-mail-1.sun.com ([192.18.19.6])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ400L1MPK6EP00@nwk-avmta-1.sfbay.Sun.COM> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Thu,
 10 Apr 2008 14:40:55 -0700 (PDT)
Received: from fe-apac-06.sun.com
 (fe-apac-06.sun.com [192.18.19.177] (may be forged))
	by sineb-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m3ALf6jF014062	for
 <psarc-ext@sun.com>; Thu, 10 Apr 2008 21:41:06 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JZ400E01PCRD300@mail-apac.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Fri,
 11 Apr 2008 05:40:19 +0800 (SGT)
Received: from [129.146.106.55] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JZ400BPSPJ5RFDH@mail-apac.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Fri, 11 Apr 2008 05:40:19 +0800 (SGT)
Date: Thu, 10 Apr 2008 14:40:51 -0700
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <18430.19272.702199.86531@gargle.gargle.HOWL>
Sender: Darren.Reed@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: Zhijun <Zhijun.Fu@sun.com>, PSARC-EXT <PSARC-ext@sun.com>
Message-id: <47FE8963.9060006@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-Accept-Language: en-au, en
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
 <18428.49039.503949.2406@gargle.gargle.HOWL> <47FD7FE4.8020601@Sun.COM>
 <18430.3958.401815.539129@gargle.gargle.HOWL> <47FE42FE.3090308@Sun.COM>
 <18430.19272.702199.86531@gargle.gargle.HOWL>
User-Agent: Mozilla/5.0 (X11; U; SunOS i86pc; en-US; rv:1.7) Gecko/20060120
Status: RO
Content-Length: 1443

James Carlson wrote:

> ...
>
>>>There may be such cases, but that's exactly what the existing "group"
>>>feature in IP filter is designed to handle.  It allows you to specify
>>>that packets matching certain criteria are to be further processed by
>>>a subset of the rules.
>>>  
>>>      
>>>
>>Right. And we're using "group" together with "ip-head".
>>    
>>
>>>This new "ip-head" and "ip-nat" feature does the same thing, 
>>>      
>>>
>>"ip-head" does the similar thing for "head", and "ip-nat" is different.
>>Please note the group feature is only available for IP Filtering rules 
>>today,  there're no group support for IP NAT rules. So "ip-nat" is 
>>irrelevant with the "group" feature.
>>    
>>
>
>That's perhaps a bug, but irrelevant.
>
>You still haven't answered what "ip-head" does that "head" doesn't
>already do.
>  
>

The difference is "ip-head" says "now go and make the packet look
like it is an IP packet (move b_rptr) and process it like an IP packet
in IPFilter starting with the specified group."  The "head" just says
go look at this set of rules next.

So, we could write rules like this:

pass in quick on bge0 family ether from any to any type 0x800 head 800
block in family ether from any to 80:00:00:00:00:00/1 group 800 ip-head 20

So we would have IP packets coming in on bge0 passed to group 800 where
those that are not broadcast/multicast are passed off for IP checking 
starting
with group 20.

Darren


From Darren.Reed@sun.com Thu Apr 10 18:23:39 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m3B1Nc8q024231
	for <psarc-ext@sac.sfbay.Sun.COM>; Thu, 10 Apr 2008 18:23:38 -0700 (PDT)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id m3B1NVbR019108
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Fri, 11 Apr 2008 09:23:37 +0800 (SGT)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ400H0DZVAZ200@brm-avmta-1.central.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Thu, 10 Apr 2008 19:23:34 -0600 (MDT)
Received: from sineb-mail-1.sun.com ([192.18.19.6])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ400DO5ZV8DZ20@brm-avmta-1.central.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Thu,
 10 Apr 2008 19:23:33 -0600 (MDT)
Received: from fe-apac-06.sun.com
 (fe-apac-06.sun.com [192.18.19.177] (may be forged))
	by sineb-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m3B1NiMc021637	for
 <psarc-ext@sun.com>; Fri, 11 Apr 2008 01:23:45 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JZ400301ZS85700@mail-apac.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Fri,
 11 Apr 2008 09:22:57 +0800 (SGT)
Received: from [129.146.106.55] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JZ400B45ZTDRF2I@mail-apac.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Fri, 11 Apr 2008 09:22:27 +0800 (SGT)
Date: Thu, 10 Apr 2008 18:22:59 -0700
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <18428.49039.503949.2406@gargle.gargle.HOWL>
Sender: Darren.Reed@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: PSARC-EXT <psarc-ext@sun.com>, Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <47FEBD73.50806@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-Accept-Language: en-au, en
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
 <18428.49039.503949.2406@gargle.gargle.HOWL>
User-Agent: Mozilla/5.0 (X11; U; SunOS i86pc; en-US; rv:1.7) Gecko/20060120
Status: RO
Content-Length: 3976

James Carlson wrote:

>Darren Reed writes:
>  
>
>>This case will extend PSARC/2005/334, by adding the ability to intercept
>>packets in MAC layer using the PFHooks infrastructure.
>>    
>>
>
>Very minor nit: case title doesn't quite match the contents.  I was
>excited to see the name, because I'm working on MAC layer interception
>... until I read that it was just a PFHooks extension for layer 2.
>  
>

What were you expecting that you didn't find?

I'm curious...


>>Users can use ipf(1M) to add ethernet filtering rules in addition to IP 
>>filtering rules, the ethernet filtering rules are marked with "family ether".
>>Unlike IPv6, no special command line switch is required to load ethernet 
>>rules. And by default, ethernet rules should be put in /etc/ipf/ipf.conf.
>>    
>>
>
>That seems strange.
>
>We currently have /etc/ipf/ipf.conf for IPv4 and the undocumented
>/etc/ipf/ipf6.conf for IPv6.  Why wouldn't we have /etc/ipf/ipfl2.conf
>(or some such) for L2-specific rules?
>
>Or if "family ether" is a good way to do this, why wouldn't we have
>"family inet" and "family inet6" and get rid of /etc/ipf/ipf6.conf?
>
>What's the intended direction?
>  
>

Eventually, the use of ipf.conf and ipf6.conf can be merged,
but the operation of merging the configurations isn't as simple as:
cat ipf.conf ipf6.conf

Use of sed to insert "family inet" and "family inet6" would be required.

There are some rules, such as these:
pass in all
block out all

that carry no protocol or version information in them.
Thus in a rule set that says this:

block in all
pass in on bge0 family ether from any to 12:34:56:78:9a:bc
pass in on bge0 from any to 1.2.3.4

the "block in all" is applied to both IPv4 and ethernet traffic.


>>The layer 2 filtering functionality is disabled by default, to enable it,
>>add the following line to the top of ipf.conf:
>>
>>set intercept_layer2 true;
>>    
>>
>
>If I have to say "family ether" in order to specify a rule that
>filters L2 packets, why do I need to give this extra command?  Doesn't
>the existence of at least one "family ether" rule mean that I intend
>to filter L2 packets as well (and thus I want interception turned on)?
>  
>

Because there are rules that will instantly apply to ethernet (see above),
it is highly likely that there will be an adverse impact on networking if
it was just enabled automatically.  Although the suggestion about
turning it on automagically if there are "family ether" rules present
does hold some water.


>>To distinguish IP filter/NAT rules intended to be processed in layer 2
>>from the rest of ipfilter rules, an additional keyword "layer2" is added.
>>Those ipfilter rules to be processed in layer 2 are marked with "layer2",
>>so these rules won't be processed again when packets goes up to IP.
>>    
>>
>
>If I have rules that have "layer2" set, then why do I need to specify
>"ip-head" or "ip-nat" in the "family ether" filter?  Shouldn't any
>rules with "layer2" set just _automatically_ match?
>  
>

There is a bunch of extra work required to make the packet
suitable for filtering against IP rules and the "ip-head" tells
ipfilter when to do that work and start doing some layer 3
comparisons.

>Or perhaps the question is this: why would I want to have rules
>specified as "layer2", but then specifically avoid sending some
>packets through those rules with "ip-head" or "ip-nat"?  If I did have
>such a case, why wouldn't I set up a "family ether" rule that
>specifies "quick" -- so that the rest of the "layer2"-tagged rules
>aren't examined at all?  That (using "quick" instead for the reverse
>sense) seems a lot clearer to me than "ip-head" or "ip-nat".
>  
>

Again, it is a function of joining the filtering of two different
layers together.  I don't think it is very clear if the ruleset
goes:

pass in on bge0 family ether ....
pass in on bge0 proto tcp ...

and for ipfilter to magically morph the ethernet packet into
an IP packet between those rules.

Darren


From Darren.Reed@sun.com Thu Apr 10 19:00:28 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m3B20RNd024667
	for <psarc-ext@sac.sfbay.Sun.COM>; Thu, 10 Apr 2008 19:00:27 -0700 (PDT)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id m3B20Mf1002239
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Fri, 11 Apr 2008 10:00:26 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ500M011KNWN00@nwk-avmta-1.sfbay.Sun.COM> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Thu, 10 Apr 2008 19:00:23 -0700 (PDT)
Received: from sineb-mail-2.sun.com ([192.18.19.7])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ500KEK1KLQY10@nwk-avmta-1.sfbay.Sun.COM> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Thu,
 10 Apr 2008 19:00:22 -0700 (PDT)
Received: from fe-apac-05.sun.com
 (fe-apac-05.sun.com [192.18.19.176] (may be forged))
	by sineb-mail-2.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m3B20pPQ016828	for
 <psarc-ext@sun.com>; Fri, 11 Apr 2008 02:00:51 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JZ500A011IFR800@mail-apac.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Fri,
 11 Apr 2008 10:00:14 +0800 (SGT)
Received: from [129.146.106.55] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JZ5004EY1KCPGTP@mail-apac.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Fri, 11 Apr 2008 10:00:14 +0800 (SGT)
Date: Thu, 10 Apr 2008 19:00:19 -0700
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <18430.3958.401815.539129@gargle.gargle.HOWL>
Sender: Darren.Reed@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: Zhijun.Fu@sun.com, PSARC-EXT <psarc-ext@sun.com>
Message-id: <47FEC633.4020502@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=us-ascii
Content-transfer-encoding: 7BIT
X-Accept-Language: en-au, en
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
 <18428.49039.503949.2406@gargle.gargle.HOWL> <47FD7FE4.8020601@Sun.COM>
 <18430.3958.401815.539129@gargle.gargle.HOWL>
User-Agent: Mozilla/5.0 (X11; U; SunOS i86pc; en-US; rv:1.7) Gecko/20060120
Status: RO
Content-Length: 855

James Carlson wrote:

> ...
>
>Yep; understood.  You'd end up with separate command line flags if you
>did that, just as IPv6 does today.
>
>My point in saying this is that the design looks incomplete.  It looks
>like you've added L2 (Ethernet, really; nothing else seems to be
>supported) to a single configuration file, but the equivalent work for
>IPv6 hasn't been done yet, so the design is "mixed" and confusing as a
>result.
>  
>

Unless I'm mistaken, it will cover two out of the three supported
MACs: ether and wifi (the odd man out being ib.)

The "family" keyword will be reused when it comes time to do the
merging of the two protocols and it is necessary to allow address-less
rules to map.

>All of the L2 examples you've provided use IPv4.  What would happen if
>I needed to add L2+IPv6 rules?  Can I do that?
>  
>

Yes, you can.

Darren


From Darren.Reed@sun.com Fri Apr 11 13:12:14 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id m3BKCDig023598
	for <psarc-ext@sac.sfbay.sun.com>; Fri, 11 Apr 2008 13:12:13 -0700 (PDT)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id m3BKC2Mw002523
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Fri, 11 Apr 2008 21:12:12 +0100 (BST)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0JZ600L19G482Z00@nwk-avmta-2.sfbay.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Fri, 11 Apr 2008 13:12:08 -0700 (PDT)
Received: from sineb-mail-1.sun.com ([192.18.19.6])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0JZ600I5BG47PY30@nwk-avmta-2.sfbay.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Fri,
 11 Apr 2008 13:12:08 -0700 (PDT)
Received: from fe-apac-05.sun.com
 (fe-apac-05.sun.com [192.18.19.176] (may be forged))
	by sineb-mail-1.sun.com (8.13.6+Sun/8.12.9) with ESMTP id m3BKCJOg016525	for
 <psarc-ext@sun.com>; Fri, 11 Apr 2008 20:12:19 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0JZ600F01G2YTV00@mail-apac.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Sat,
 12 Apr 2008 04:11:59 +0800 (SGT)
Received: from [129.146.106.55] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0JZ60049PG3YPGGV@mail-apac.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Sat, 12 Apr 2008 04:11:59 +0800 (SGT)
Date: Fri, 11 Apr 2008 13:12:04 -0700
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <18428.49039.503949.2406@gargle.gargle.HOWL>
Sender: Darren.Reed@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: PSARC-EXT <psarc-ext@sun.com>, Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <47FFC614.6030902@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=us-ascii
Content-transfer-encoding: 7BIT
X-Accept-Language: en-au, en
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
 <18428.49039.503949.2406@gargle.gargle.HOWL>
User-Agent: Mozilla/5.0 (X11; U; SunOS i86pc; en-US; rv:1.7) Gecko/20060120
Status: RO
Content-Length: 509

Jim,

I thought about some of your comments, especially those about it
being very ethernet centric and I've changed it to "waiting need spec"
as the team needs to rethink what it is they're designing and putting
forward here as there are some serious problems with it (that haven't
yet been raised) as it stands today.

In addition, there is clearly room for more descriptive text concerning
how the changes to ipf.conf work.

We should move further discussion on this to networking-discuss.

Thanks,
Darren


From Darren.Reed@sun.com Sat Nov 22 02:38:20 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mAMAc4SM011222
	for <psarc-ext@sac.sfbay.sun.com>; Sat, 22 Nov 2008 02:38:20 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id mAMAbnhM008095
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Sat, 22 Nov 2008 10:37:53 GMT
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KAQ0090TDJ19000@brm-avmta-1.central.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Sat, 22 Nov 2008 03:37:49 -0700 (MST)
Received: from gmp-eb-inf-1.sun.com ([192.18.6.21])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KAQ00H6FDIZXK50@brm-avmta-1.central.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Sat,
 22 Nov 2008 03:37:48 -0700 (MST)
Received: from fe-emea-09.sun.com (gmp-eb-lb-1-fe3.eu.sun.com [192.18.6.10])
	by gmp-eb-inf-1.sun.com (8.13.7+Sun/8.12.9) with ESMTP id mAMAbl3f019945	for
 <psarc-ext@sun.com>; Sat, 22 Nov 2008 10:37:47 +0000 (GMT)
Received: from conversion-daemon.fe-emea-09.sun.com by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KAQ00C01DEPG600@fe-emea-09.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Sat,
 22 Nov 2008 10:37:47 +0000 (GMT)
Received: from [129.157.21.188] by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KAQ009WMDIYCG90@fe-emea-09.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Sat, 22 Nov 2008 10:37:47 +0000 (GMT)
Date: Sat, 22 Nov 2008 02:37:44 -0800
From: Darren Reed <Darren.Reed@sun.com>
Subject: PSARC/2008/249 - Packet Interception for the MAC layer
Sender: Darren.Reed@sun.com
To: psarc-ext@sun.com
Cc: Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <4927E0F8.8080103@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
User-Agent: Thunderbird 2.0.0.18 (Windows/20081105)
Status: RO
Content-Length: 17757

After an extended period on the sidelines, the
layer 2 filtering project is resubmitting this
PSARC fast track for consideration. No changes
are being provided between this and the original
specification as various parts of the design have
been completely revisited, effectively obsoleting
the original specification.

Additional materials (diff marked man pages)
can be found in the case directory:
gld.7d.txt
hook_nic_event.9s.txt
ipf.4.txt
ipnat.4.txt
net_getlifaddr.9f.txt

Darren


Abstract
========
This case will extend PSARC/2005/334, by adding the ability to intercept
packets in MAC layer using the PFHooks infrastructure.

This case only makes one change, an addition, to the interfaces that were
committed to by PSARC/2008/219 (see "new NIC event" below for details)

Release Biding
--------------
This case seeks for a patch binding.

Introduction
============
The PFHooks project, PSARC/2005/334, provide the ability to intercept 
packets
in IP layer by adding hooks into the network stack.

Since its integration, there has been customer requirements for the ability
to intercept packets in MAC layer as well, also it is needed to enforce
security rules for xVM guest domains and exclusive zones.

Goals
-----
This case seeks to meet the following goals:
* provide the hooks in MAC layer that allows consumers to register on to
  intercept packets;

* provide the netinfo interface for MAC layer that gives consumers access to
  interface information, and the ability to inject or emit packets directly;

* modify IPFilter to allow the administrator to specify layer 2 rules, which
  includes ethernet filtering rules and IP Filtering/NAT rules.

Out of scope
------------
This project only provides the ability to specify ethernet filtering rules
to match ethernet packets, and IP filtering/NAT rules to match/modify IP
packets at MAC layer. Providing the ability to specify rules to filter
non-ethernet packets by matching the MAC header is out of scope for this
project.

The detailed design are described below for each major components.

Netinfo interface for MAC layer
===============================
netinfo interfaces
------------------
The hooks provided will generate events for NH_PHYSICAL_IN and 
NH_PHYSICAL_OUT,
using the same interface as IPv4 and IPv6 do in PSARC/2005/334.

The following functions will be supported through the netinfo(9f) framework:
net_getifname()
net_phylookup()
net_phygetnext()
net_getlifaddr()
net_inject()
net_getmtu()

All of the other functions in the netinfo(9f) framework will return a value
indicating that they are unsupported. The return values for the above
functions only have meaning with the scope of the corresponding family -
it is not correct to use a value returned by net_getifname() using the
ethernet net_data_t handle with net_phylookup() for IPv4.

The callback for these events will receive a pointer to a hook_pkt_event_t
structure that has the following fields filled out:

hpe_ifp - 0 for NH_PHYSICAL_OUT, otherwise a value indicating which
          interface the NH_PHYSICAL_IN event is associated with;
hpe_ofp - 0 for NH_PHYSICAL_IN, otherwise a value indicating which
          interface the NH_PHYSICAL_OUT event is associated with;
hpe_hdr - points to the start of the MAC header
hpe_mb  - points to the start of the mblk_t that holds hpe_hdr;
hpe_mp  - points to the mblk_t that is the start of the packet.

Name to interface resolution
----------------------------
After Clearview UV all the data link related operations use link names, this
applies to IPFilter as well. When the administrator wants to specify a rule
that works on certain interface, link name is used to specify which 
interface
this rule applies to. So link name consititutes the interface name for
MAC layer netinfo.

Since layer 2 filtering is based on the MAC client which Crossbow project is
introducing, in this project we'll introduce the MAC client index as the
MAC layer interface pointer, to uniquely indentify a MAC layer interface
in the kernel. This is similar to the existing ifindex that is used as IP
layer interface pointer today.

Netinfo provides functions to translate from a interface name (link name)
to the corresponding interface pointer (MAC client index) and back, via
net_phylookup() and net_getifname(). And these functions can be called in
data path so the existing procedures such as dls_mgmt_get_linkid() and
dls_mgmt_get_linkinfo() cannot be used as they involve door calls.
Thus we propose to add a link name <-> link id hash table in dls, and 
provide
the following routines to translate between link name and mac name. The MAC
layer netinfo will use these routines to implement mapping between link name
and MAC client index.

+------------------------------------------------------------+
| Interface                                 | Classification |
|------------------------------------------------------------|
| dls_devnet_macname2linkname(const char *, |                |
|     char *, const size_t);                | consolidation  |
| dls_devnet_linkname2macname(const char *, | private        |
|     char *, const size_t);                |                |
+------------------------------------------------------------+
     Table: Fuctions for link name and mac name mapping

new NIC event
-------------
The status of network in the operating system often changes, from unplugging
a system from network temporarily, to an interface's IP address changing
as a result of DHCP. Thus PFHooks framework provides event notification
mechanism for this.

The callback for these events will receive a pointer to a hook_nic_event_t
structure that has the following fields filled out:

hne_protocol - network protocol for events, returned from net_lookup

hne_nic      - physical interface associated with event

hne_lif      - logical interface (if any) associated with event

hne_event    - type of event occuring. The current list of events 
available is:

        NE_PLUMB
               indicates that an interface has just been created

        NE_UNPLUMB
               indicates that an interface has just been destroyed and that
               no more events should be received for it

        NE_UP
               indicates that an interface has changed state to "up" and
               may now generate packet events.

        NE_DOWN
               indicates that an interface has changed state to "down" and
               will no longer generate packet events.

        NE_ADDRESS_CHANGE
               indicates that an address on an interface has changed.

hne_data     - pointer to extra data about event or NULL if none

hne_datalen  - size of data pointed to by hne_data (can be 0)

NE_NAME_CHANGE event
~~~~~~~~~~~~~~~~~~~~
As Clearview UV (PSARC/2006/499, PSARC/2007/527, PSARC/2008/002) introduces
the ability to rename a data link, we need to capture this event in order to
update IPFilter rules accrodingly. Thus we propose an extension to
PSARC/2008/219 by adding a new hook event NE_NAME_CHANGE to nic_event_t
to indicate the that an interface has been renamed, and this particular 
event
is only available to layer 2 netinfo. In IP, changing of an interface name
is represented by a NE_UNPLUMB and NE_PLUMB event pair.

typedef enum nic_event {
         NE_PLUMB = 1,
         NE_UNPLUMB,
         NE_UP,
         NE_DOWN,
         NE_ADDRESS_CHANGE,
+        NE_NAME_CHANGE
} nic_event_t;

Design considerations
~~~~~~~~~~~~~~~~~~~~~
IPFilter rules always match by name, and only the current link names are 
used
for matching, not old names. Uppon NE_NAME_CHANGE event, IPFilter will walk
all the layer 2 rules, and resolve the interface name stored in the rule
structure into interface pointers. So when the link is renamed, rules using
old link names are invalidated, and rules using new link names are 
activated.
If there's a filtering rule that applies to interface bge0, and someone 
renames
bge0 to net0, then the rule no longer matches packets received on the link
formally known as bge0.

Also IPFilter has been designed to allow users to specify rules with 
interface
names that do not exist at the time they are loaded, and for those interface
names to be resolved at the time at which they're added to the system. Thus,
the mapping from the linkname to the linkid needs to happen in the kernel.
Changing IPFilter to use linkid instead of link name will not work.

Protocol & Hook registration
============================
Protocol registration
---------------------
With IP layer netinfo today we have 3 protocols, IPv4, IPv6 and ARP. For MAC
layer, each of the MAC plugin type is treated as a different protocol, so
we'll have ethernet, wifi and ib. These protocols will be registered by
using net_protocol_register() when the corresponding MAC plugin gets loaded.

Hook registration
-----------------
IPFilter will register hooks for MAC layer protocols in the following cases:

when the first ethernet filtering rule is added
- register the ethernet hook

when the first "layer2" IP filtering/NAT rule is added
- register the ethernet, wifi and ib hooks

when it receives a notification indicating that a protocol is registered
- register the hook if there are rules for that corresponding protocol.

Since layer 2 filtering functionality is enabled automatically when the
first layer 2 rule is added, the corresponding hook needs to be registered
then so packets can be passed to IPFilter from the hook framework.

It is possible that a rule for a layer 2 protocol is added before the
corresponding protocol is registered. Suppose user has added a layer 2 IP
filtering rule on a system that only has ethernet cards, then he plugs a
wifi card into the system and sets it up, in this case when the wifi MAC
plugin is loaded, the protocol will be registered, and IPFilter will be
notified via the callback notification mechanism provided by the PFHooks
API project, and it will register the hook for that protocol so it can
receive and match wifi packets.

IPFilter changes
================
Users can use ipf(1M) to add ethernet filtering rules in addition to IP
filtering rules, these ethernet filtering rules are marked with "family 
ether".
They can also add IP Filtering/NAT rules and mark them with "layer2" keyword
so these rules will be processed in MAC layer instead of IP layer. 
Unlike IPv6,
no special command line switch is required to load these rules.

The "layer2" IP filtering/NAT rules go to existing ipf.conf, ipf6.conf and
ipnat.conf, respectively. The "family ether" rules go to a new configuration
file ipf-ether.conf.

The layer 2 filtering functionality will be enabled automatically when the
first ethernet rule or "layer2" IPFilter rule is added, and disabled when
the last such rule is removed. This functionality is only available in
global zone.

Also, ipmon has been updated to print out log records with ethernet
information but the output of this command is volatile.

Rule processing
---------------
Currently processing order in IPFilter is:

[INPUT] -> IP NAT -> IP firewall -> { IP }  -> IP firewall -> IP NAT -> 
[OUTPUT]

With layer 2 filtering the processing order would become:

[INPUT] -> L2 firewall -> "layer2" IP NAT -> "layer2" IP firewall ->
... -> IP NAT -> IP firewall -> { IP }  -> IP firewall -> IP NAT -> ...
-> L2 firewall -> "layer2" IP firewall -> "layer2" IP NAT -> [OUTPUT]

Input processing
~~~~~~~~~~~~~~~~
Take input processing for an IP packet for example:

- MAC level filtering rules are processed first. These rules match on MAC
headers to determine if a packet should be passed or blocked. Administrators
use these rules to match with MAC addresses, MAC type, VLAN ID, .etc.

- L2filter jump over the MAC header, determine if this is an IP packet, and
do some sanity checking before passing it up to "layer2" IP rules for 
further
processing.

- Then "layer2" IP NAT rules are processed. Like IP layer NAT rules, these
rules do NAT for IP packets, but it is done at MAC layer instead of IP 
layer.

- Then "layer2" IP Filtering rules are processed. These rules provide IP
Filtering at MAC layer.

- L2filter finishes processing and the packet is delivered up in the stack.
When the packet reaches IP, IP layer filtering/NAT processing is invoked,
and it works just as it does today.

Design considerations
~~~~~~~~~~~~~~~~~~~~~
This processing order is designed so that

- The processing order between IP NAT rules and IP Filtering rules is
consistent with existing IPFilter today;

- Since MAC level filtering rules is processed before layer2 IP rules,
down the road it is possible to combine filtering at both level together,
allow or block a packet based on a mixture of L2 and L3 criteria, thus
providing more fine grained control.

Changes to output
-----------------
With layer 2 filtering, each type of rules have its own distinct orders,
the output of ipfstat/ipnat has been modified so that the rules are shown
in a manner to let the users better understand the processing orders.
The change only applies to global zone, output in non-global zones remain
unchanged.

Example
~~~~~~~

# ipfstat -io
Ethernet rules:
empty list for ipfilter(out)
pass in family ether all
pass in family ether from 1:2:3:4:5:6 to any

layer 2 IP rules:
empty list for ipfilter(out)
pass in proto icmp from 1.1.1.1 to 2.2.2.2 layer2
block in proto tcp from 3.3.3.3 to 4.4.4.4 layer2

IP rules:
pass in all
pass out all

# ipnat -l
List of layer 2 active MAP/Redirect filters:
map bge1 from 2.3.4.5/32 to 6.7.8.9/32 -> 1.1.2.2/32 layer2

List of active MAP/Redirect filters:

List of active sessions:

Examples
--------
Prevent MAC address spoofing
~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Suppose we have a domU with a interface vnic0, we may want to ensure:
- packets from this domU can have use its own source MAC address, preventing
this domU from pretending someone else
- packet from this source MAC address can only come from this domU, 
preventing
others from pretending this domU

say vnic0 has MAC address 11:22:33:44:55:66, the rules would be 
something like:

block out family ether from 11:22:33:44:55:66 to any
block out on vnic0 family ether from any to any
pass out on vnic0 family ether from 11:22:33:44:55:66 to any

Prevent IP address spoofing
~~~~~~~~~~~~~~~~~~~~~~~~~~~

Suppose we'd want to prevent a domU from using others' IP addresses, we can
probably go with:

block out on vnic0 from any to any layer2
pass out on vnic0 from 1.1.1.1 to any layer2

while 1.1.1.1 is the assigned IP address on vnic0

VLAN packets filtering
~~~~~~~~~~~~~~~~~~~~~~

Say we'd want to block some IP traffic, below are examples on how it is done
with regard to vlan:

- block all IP traffic regardless of VLAN

block in family ether type 0x800

- block all IP traffic belonging to VLAN:

block in family ether type 0x800 with vlan

- block all IP traffic NOT belonging to VLAN:

block in family ether type 0x800 with not vlan

- block all IP traffic for a specific VLAN (e.g. 100)

block in family ether type 0x800 vlan 100

Ioctl compatibility
-------------------
ABI compatibility with the old structure definitions is preserved by 
this case.

IPFILTER_VERSION (see ipnat(7i)) is used to keep track of user application's
version thus the old binaries can still work after this change. The 
kernel code
would handle the ioctl input/output based on the version number to make it a
compatible change. There's no change required for user applications 
using the
interfaces.

The related data structures natlookup_t and nat_t remain the same, and 
ioctls
SIOCGNATL/SIOCSTPUT will work correctly. User can set a flag, IPN_LAYER2,
in natlookup_t and nat_t, respectively, to indicate it is looking 
up/inserting
a layer 2 NAT session, or a layer 3 one. For compatibilities, by default the
flag is not set, which indicates a layer 3 session.


Interfaces
==========
+----------------------------------------+----------------+
| Interface                              | Classification |
+----------------------------------------+----------------+
| dls_devnet_macname2linkname            |     Private    |
| dls_devnet_linkname2macname            |     Private    |
+----------------------------------------+----------------+
| NE_NAME_CHANGE                         |    Committed   |
| NHF_ETHER                              |    Committed   |
| NHF_WIFI                               |    Committed   |
| NHF_IB                                 |    Committed   |
+----------------------------------------+----------------+
| "ipfilter_hook_eth_in"                 |   Uncommitted  |
| "ipfilter_hook_eth_out"                |   Uncommitted  |
| "ipfilter_hook_wifi_in"                |   Uncommitted  |
| "ipfilter_hook_wifi_out"               |   Uncommitted  |
| "ipfilter_hook_ib_in"                  |   Uncommitted  |
| "ipfilter_hook_ib_out"                 |   Uncommitted  |
+----------------------------------------+----------------+
| "family ether"                         |     Volatile   |
| "layer2"                               |     Volatile   |
+----------------------------------------+----------------+
| IPFILTER_VERSION                       |    Committed   |
| ioctl SIOCGNATL                        |    Committed   |
| ioctl SIOCSTPUT                        |    Committed   |
| ioctl SIOCSTLCK                        |    Committed   |
| struct natlookup                       |   Uncommitted  |
| struct nat                             |   Uncommitted  |
| IPN_LAYER2                             |     Volatile   |
| /usr/include/netinet/ipl.h             |   Uncommitted  |
| /usr/include/netinet/ip_fil.h          |   Uncommitted  |
| /usr/include/netinet/ip_nat.h          |   Uncommitted  |
+----------------------------------------+----------------+


From Darren.Reed@sun.com Wed Dec  3 21:08:04 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mB4584U9027257
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 3 Dec 2008 21:08:04 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mB457thr046142
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Wed, 3 Dec 2008 22:08:03 -0700 (MST)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBC0040F69DML00@nwk-avmta-1.sfbay.Sun.COM> for psarc-ext@sun.com
 (ORCPT psarc-ext@Sun.COM); Wed, 03 Dec 2008 21:08:01 -0800 (PST)
Received: from gmp-eb-inf-1.sun.com ([192.18.6.21])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBC00E2069BMW70@nwk-avmta-1.sfbay.Sun.COM> for
 psarc-ext@sun.com (ORCPT psarc-ext@Sun.COM); Wed,
 03 Dec 2008 21:08:00 -0800 (PST)
Received: from fe-emea-09.sun.com (gmp-eb-lb-1-fe3.eu.sun.com [192.18.6.10])
	by gmp-eb-inf-1.sun.com (8.13.7+Sun/8.12.9) with ESMTP id mB457xt8017991	for
 <psarc-ext@Sun.COM>; Thu, 04 Dec 2008 05:07:59 +0000 (GMT)
Received: from conversion-daemon.fe-emea-09.sun.com by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KBC00301647T900@fe-emea-09.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@Sun.COM (ORCPT psarc-ext@Sun.COM); Thu,
 04 Dec 2008 05:07:59 +0000 (GMT)
Received: from [129.158.90.153] by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KBC00H62699T4C0@fe-emea-09.sun.com> for psarc-ext@Sun.COM
 (ORCPT psarc-ext@Sun.COM); Thu, 04 Dec 2008 05:07:59 +0000 (GMT)
Date: Thu, 04 Dec 2008 16:07:53 +1100
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <4927E0F8.8080103@Sun.COM>
Sender: Darren.Reed@sun.com
To: psarc-ext@sun.com
Cc: Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <493765A9.7080103@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM>
User-Agent: Thunderbird 2.0.0.18 (Windows/20081105)
Status: RO
Content-Length: 6880

With the extension of the timer for this project to tbe 10th of December,
I'm updating the materials submmitted to cover changes required to
support the eventual filtering of WiFi packets.

The changes are (see below for more details):
- additional section in "spec.txt" covering the reason, why for, etc.
- changes to hook_pkt_event(9s)

Darren


--- spec.txt    Wed Dec  3 11:30:31 2008
+++ spec.txt.new    Wed Dec  3 13:37:48 2008
@@ -3,8 +3,9 @@
 This case will extend PSARC/2005/334, by adding the ability to intercept
 packets in MAC layer using the PFHooks infrastructure.
 
-This case only makes one change, an addition, to the interfaces that were
-committed to by PSARC/2008/219 (see "new NIC event" below for details)
+This case only makes two changes, an addition, to the interfaces that were
+committed to by PSARC/2008/219 (see "hook_pkt_event_t" and "new NIC event"
+below for details)
 
 Release Biding
 --------------
@@ -72,7 +73,42 @@
 hpe_hdr - points to the start of the MAC header
 hpe_mb  - points to the start of the mblk_t that holds hpe_hdr;
 hpe_mp  - points to the mblk_t that is the start of the packet.
+hpe_hpeinfo - points to mac_header_info_t which contains MAC header 
information
 
+hook_pkt_event_t
+----------------
+In order to intercept IP packets at MAC layer, IPFilter needs to know the
+size of the MAC header to locate the IP header start. The problems is 
the wifi
+header is not self explained, parsing it requires information from mac 
handle
+thus IPFilter cannot do the parsing itself, so we need to rely on the MAC
+plugin to parse the header, pass the information through Hook framework to
+IPFilter, so it can identify IP header start correctly.
+
+While adding a header length field to hook_pkt_info_t solves the 
problem above,
+down the road we may want to provide the ability to match wifi header, 
which
+requires information of the wifi header fields in IPFilter, not just 
the header
+length, thus we propose to add a pointer to hook_pkt_info_t, which 
points at
+a structure of mac_header_info_t, and pass this through the Hook framework,
+so the hook consumers, like IPFilter, can have the needed information for
+the MAC header. The new hook_pkt_info_t strucuture would look like:
+
+typedef struct hook_pkt_event {
+        net_handle_t            hpe_protocol;
+        phy_if_t                hpe_ifp;
+        phy_if_t                hpe_ofp;
+        void                    *hpe_hdr;
+        mblk_t                  **hpe_mp;
+        mblk_t                  *hpe_mb;
+        int                     hpe_flags;
+-       void                    *hpe_reserved[2];
++       void                    *hpe_hdrinfo;
++       void                    *hpe_reserved[1];
+} hook_pkt_event_t;
+
+For existing IP/ARP Hooks, the header format is self explained, so 
hpe_hdrinfo
+will be NULL and IPFilter does the header parsing itself as before.
+
+
 Name to interface resolution
 ----------------------------
 After Clearview UV all the data link related operations use link names, 
this





Kernel functions			       hook_pkt_event(9s)

NAME
     hook_pkt_event - packet event structure  passed  through  to
     hooks

SYNOPSIS
     #include <sys/neti.h>
     #include <sys/hook.h>
     #include <sys/hook_event.h>

INTERFACE LEVEL
     Solaris DDI specific (Solaris DDI)

DESCRIPTION
     hook_pkt_event contains fields that relate	to a packet as is
     in	 a  network  protocol  handler.	 This structure	is passed
     through to	a  callback  for  NH_PRE_ROUTING,  NH_FORWARDING,
     NH_POST_ROUTING, NH_LOOPBACK_IN and NH_LOOPBACK_OUT events.

     A callback	may modify the hpe_hdr,	hpe_mp and hpe_mb fields.
     A callback	may not	modify either of hpe_ifp or hpe_ofp.

     The following table documents which  fields  can  be  safely
     used as a result of each event.
      _________________________________________________________________
     |	   Event       | hpe_ifp | hpe_ofp | hpe_hdr | hpe_mp |	hpe_mb |
     |_________________|_________|_________|_________|________|________|
     | NH_PRE_ROUTING  |   Yes	 |	   |   Yes   |	 Yes  |	  Yes  |
     | NH_POST_ROUTING |	 |   Yes   |   Yes   |	 Yes  |	  Yes  |
     | NH_FORWARDING   |   Yes	 |   Yes   |   Yes   |	 Yes  |	  Yes  |
     | NH_LOOPBACK_IN  |   Yes	 |	   |   Yes   |	 Yes  |	  Yes  |
     | NH_LOOPBACK_OUT |	 |   Yes   |   Yes   |	 Yes  |	  Yes  |
     |_________________|_________|_________|_________|________|________|

STRUCTURE MEMBERS
     net_data_t	  hpe_family;
     phy_if_t	  hpe_ifp;
     phy_if_t	  hpe_ofp;
     void	  *hpe_hdr;
     mblk_t	  **hpe_mp;
     mblk_t	  *hpe_mb;
     uint32_t	  hpe_flags;
+    void	  *hpe_hdrinfo;

     hpe_family
	  The protocol family for this packet.	The  value  found
	  here will match the corresponding value returned from	a
	  call to net_protocol_lookup(9f).

     hpe_ifp
	  The inbound interface	for a packet.

     hpe_ofp
	  The outbound interface for a packet.

     hpe_hdr
	  Pointer to the start of the  network	protocol  header,
	  within an mblk_t.

     hpe_mp
	  Pointer to the mblk_t	pointer	that points to the  first
	  mblk_t in this packet.

     hpe_mb
	  Pointer to the mblk_t	that contains hpe_hdr.

+    hpe_hdrinfo
+         Optional pointer to packet header information.

     The relationship of hpe_hdr, hpe_mp and hpe_mb can	be illus-
     trated as follows.

       hpe_mp
	 |
	 V
	mb	      hpe_mb
	 |		|
	 V		V
     +--------+	   +----------+	   +--------+ __-------->+-------+
     | mblk_t |	   | mblk_t   |	   | b_rptr--/		 |data	 |
     | b_cont----->| db_datap----->| b_wptr--\	hpe_hdr->|header |
     +--------+	   +----------+	   +--------+ \		 |data	 |
					       \	 +-------+
						\________________^

     hpe_flags
	  This field is	used to	carry  additional  properties  of
	  packets.  THe	current	collection of defined bits is:

	  HPE_BROADCAST
	       Set if the packet was recognised	 as  a	broadcast
	       packet  from  the  link	layer.	 Cannot	be set if
	       HPE_MULTICAST is	set, currently oly possible  with
	       physical	in packat events.

	  HPE_MULTICAST
	       Set if the packet was recognised	 as  a	multicast
	       packet  from  the  link	layer.	 Cannot	be set if
	       HPE_BROADCAST is	set, currently oly possible  with
	       physical	in packat events.

ATTRIBUTES
     See attributes(5) for descriptions	of the	following  attri-
     butes:
	  ____________________________________________________________
	 |	 ATTRIBUTE TYPE	       |       ATTRIBUTE VALUE	     |
	 |_____________________________|_____________________________|
	 | Interface Stability	       |	 Committed	     |
	 |_____________________________|_____________________________|

SunOS 5.10	   Last	change:	DD Month 2007			2

Kernel functions			       hook_pkt_event(9s)

SEE ALSO
     netinfo(9f),

SunOS 5.10	   Last	change:	DD Month 2007			3







From carlsonj@phorcys.east.sun.com Wed Dec 10 07:01:20 2008
Received: from sunmail3mpk.sfbay.sun.com (sunmail3mpk.SFBay.Sun.COM [129.146.11.52])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBAF1KSD024631
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 10 Dec 2008 07:01:20 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail3mpk.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mBAF17WV015421;
	Wed, 10 Dec 2008 07:01:17 -0800 (PST)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBO00B4P1Q3MT00@nwk-avmta-2.sfbay.sun.com>; Wed,
 10 Dec 2008 07:01:15 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBO0018C1PZMCD0@nwk-avmta-2.sfbay.sun.com>; Wed,
 10 Dec 2008 07:01:12 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id mBAF1BxG014225; Wed,
 10 Dec 2008 10:01:11 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id mBAF1BtZ014222; Wed,
 10 Dec 2008 10:01:11 -0500 (EST)
Date: Wed, 10 Dec 2008 10:01:11 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <493765A9.7080103@Sun.COM>
To: Darren Reed <Darren.Reed@sun.com>
Cc: psarc-ext@sun.com, Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <18751.55735.419739.64295@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
Status: RO
Content-Length: 1843

Darren Reed writes:
> With the extension of the timer for this project to tbe 10th of December,
> I'm updating the materials submmitted to cover changes required to
> support the eventual filtering of WiFi packets.

A few quick questions and checks:

1.  I assume that "net_getlifaddr" in the context of a MAC hook
    actually returns a MAC layer address, and not an IP address as the
    name "lif" might otherwise suggest.  Right?

2.  You give the new rule processing as:

    [INPUT] -> L2 firewall -> "layer2" IP NAT -> "layer2" IP firewall ->
    ... -> IP NAT -> IP firewall -> { IP }  -> IP firewall -> IP NAT -> ...
    -> L2 firewall -> "layer2" IP firewall -> "layer2" IP NAT -> [OUTPUT]

    I don't understand why the ordering of operations isn't just
    reversed on output.  Assuming the input order is correct (and it
    appears to me to be right), the "L2 firewall" element should be
    the last operation before "[OUTPUT]."

3.  We're defining these bits of syntax ourselves, and we're expecting
    that administrators are going to rely on them for the security of
    their systems.  Given that, is "Volatile" the right classification
    for the new "family ether" and "layer2" configuration keywords?

4.  Why is hpe_hdrinfo "void *" rather than "hook_pkt_info_t *"?  Void
    pointers are mildly evil, as they prevent the compiler and lint
    from doing their type-checking jobs.  What else would hpe_hdrinfo
    point to?

(One very small code review nit: as an argument in a function
definition [dls_devnet_*name2*name], 'const size_t' doesn't do
anything that 'size_t' alone wouldn't do.)

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Zhijun.Fu@sun.com Wed Dec 10 08:10:28 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBAGAR07020736
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 10 Dec 2008 08:10:28 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id mBAGAInS028319
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Wed, 10 Dec 2008 16:10:27 GMT
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBO0043L4XEY600@brm-avmta-1.central.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Wed, 10 Dec 2008 09:10:26 -0700 (MST)
Received: from sineb-mail-2.sun.com ([192.18.19.7])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBO00FLJ4XD7VB0@brm-avmta-1.central.sun.com> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Wed,
 10 Dec 2008 09:10:26 -0700 (MST)
Received: from fe-apac-05.sun.com
 (fe-apac-05.sun.com [192.18.19.176] (may be forged))
	by sineb-mail-2.sun.com (8.13.6+Sun/8.12.9) with ESMTP id mBAGAO6m014891	for
 <psarc-ext@sun.com>; Wed, 10 Dec 2008 16:10:24 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0KBO001014RQKJ00@mail-apac.sun.com> (original mail from Zhijun.Fu@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Thu,
 11 Dec 2008 00:10:24 +0800 (SGT)
Received: from [129.150.144.21] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0KBO001OL4X0GNTJ@mail-apac.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Thu, 11 Dec 2008 00:10:24 +0800 (SGT)
Date: Thu, 11 Dec 2008 00:10:11 +0800
From: Zhijun Fu <Zhijun.Fu@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <18751.55735.419739.64295@gargle.gargle.HOWL>
Sender: Zhijun.Fu@sun.com
To: James Carlson <james.d.carlson@sun.com>
Cc: Darren Reed <Darren.Reed@sun.com>, PSARC-ext@sun.com
Message-id: <493FE9E3.9000005@sun.com>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
 <18751.55735.419739.64295@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.18 (Windows/20081105)
Status: RO
Content-Length: 3115

James Carlson wrote:
> Darren Reed writes:
>> With the extension of the timer for this project to tbe 10th of December,
>> I'm updating the materials submmitted to cover changes required to
>> support the eventual filtering of WiFi packets.
>
> A few quick questions and checks:

Jim,

Thank you very much for your help and comments!

>
> 1.  I assume that "net_getlifaddr" in the context of a MAC hook
>     actually returns a MAC layer address, and not an IP address as the
>     name "lif" might otherwise suggest.  Right?

Yes, that's right. For MAC hook we're kind of reusing the same function
though there is actually no "lif" at MAC layer. :-) I'd like to hear your
suggestions on this.

>
> 2.  You give the new rule processing as:
>
>     [INPUT] -> L2 firewall -> "layer2" IP NAT -> "layer2" IP firewall ->
>     ... -> IP NAT -> IP firewall -> { IP }  -> IP firewall -> IP NAT -> ...
>     -> L2 firewall -> "layer2" IP firewall -> "layer2" IP NAT -> [OUTPUT]
>
>     I don't understand why the ordering of operations isn't just
>     reversed on output.  Assuming the input order is correct (and it
>     appears to me to be right), the "L2 firewall" element should be
>     the last operation before "[OUTPUT]."

It could be a useful feature to combine filtering on L2 and L3
together, or even trigger conditional IP NAT from a L2 filtering rule,
though they are not supported in the current project scope, down the
road we may want to add it, and these features require that L2 firewall
to be processed before those "layer2" IP rules.

Also, since the packet is received at MAC layer, it might be more
straight forward to do the parsing and matching from the MAC header,
as we need to parse the MAC header first anyway to locate the IP header
start. It also allows a simpler implementation in IPFilter but that
can be considered an implementation issue anyway.

Would the above make sense?

>
> 3.  We're defining these bits of syntax ourselves, and we're expecting
>     that administrators are going to rely on them for the security of
>     their systems.  Given that, is "Volatile" the right classification
>     for the new "family ether" and "layer2" configuration keywords?

We'll think more about this.

>
> 4.  Why is hpe_hdrinfo "void *" rather than "hook_pkt_info_t *"?  

For MAC layer, it is actually "mac_header_info_t *".

> Void
>     pointers are mildly evil, as they prevent the compiler and lint
>     from doing their type-checking jobs.  What else would hpe_hdrinfo
>     point to?

The major reason for using "void *" here, is that hook_pkt_event_t
is a quite general structure that works with all protocols, thus I
feel it is probably not appropriate to use types that specific to
one protocol, e.g. mac_header_info_t is specific to MAC. Also it
could point to other data structure for protocol other than MAC
layer, though there's not such need right now.

>
> (One very small code review nit: as an argument in a function
> definition [dls_devnet_*name2*name], 'const size_t' doesn't do
> anything that 'size_t' alone wouldn't do.)
>   

Thank you, will fix.

Regards,
Zhijun

From carlsonj@phorcys.east.sun.com Wed Dec 10 08:26:53 2008
Received: from sunmail2sca.sfbay.sun.com (sunmail2sca.SFBay.Sun.COM [129.145.155.234])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBAGQrCF013418
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 10 Dec 2008 08:26:53 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail2sca.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mBAGQmsA012167;
	Wed, 10 Dec 2008 08:26:50 -0800 (PST)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBO00F0Z5OQXD00@nwk-avmta-2.sfbay.sun.com>; Wed,
 10 Dec 2008 08:26:50 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBO00DGQ5OODL40@nwk-avmta-2.sfbay.sun.com>; Wed,
 10 Dec 2008 08:26:49 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id mBAGQmAX014882; Wed,
 10 Dec 2008 11:26:48 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id mBAGQme0014879; Wed,
 10 Dec 2008 11:26:48 -0500 (EST)
Date: Wed, 10 Dec 2008 11:26:48 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <493FE9E3.9000005@sun.com>
To: Zhijun Fu <Zhijun.Fu@sun.com>
Cc: Darren Reed <Darren.Reed@sun.com>, PSARC-ext@sun.com
Message-id: <18751.60872.537103.651790@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
 <18751.55735.419739.64295@gargle.gargle.HOWL> <493FE9E3.9000005@sun.com>
Status: RO
Content-Length: 4057

Zhijun Fu writes:
> James Carlson wrote:
> > 1.  I assume that "net_getlifaddr" in the context of a MAC hook
> >     actually returns a MAC layer address, and not an IP address as the
> >     name "lif" might otherwise suggest.  Right?
> 
> Yes, that's right. For MAC hook we're kind of reusing the same function
> though there is actually no "lif" at MAC layer. :-) I'd like to hear your
> suggestions on this.

I've no problem with this.  I was just verifying that I understood, as
it doesn't seem to be in the materials.

> > 2.  You give the new rule processing as:
[...]
> >     I don't understand why the ordering of operations isn't just
> >     reversed on output.  Assuming the input order is correct (and it
> >     appears to me to be right), the "L2 firewall" element should be
> >     the last operation before "[OUTPUT]."
> 
> It could be a useful feature to combine filtering on L2 and L3
> together, or even trigger conditional IP NAT from a L2 filtering rule,
> though they are not supported in the current project scope, down the
> road we may want to add it, and these features require that L2 firewall
> to be processed before those "layer2" IP rules.
> 
> Also, since the packet is received at MAC layer, it might be more
> straight forward to do the parsing and matching from the MAC header,
> as we need to parse the MAC header first anyway to locate the IP header
> start. It also allows a simpler implementation in IPFilter but that
> can be considered an implementation issue anyway.
> 
> Would the above make sense?

I'm not sure, but I don't think any of it answers the actual question
I had.  You currently have this:

     [INPUT] -> L2 firewall -> "layer2" IP NAT -> "layer2" IP firewall ->
    ... -> IP NAT -> IP firewall -> { IP }  -> IP firewall -> IP NAT -> ...
     -> L2 firewall -> "layer2" IP firewall -> "layer2" IP NAT -> [OUTPUT]

But that's not entirely symmetric, and I don't see why.  I would have
expected it to be:

     [INPUT] -> L2 firewall -> "layer2" IP NAT -> "layer2" IP firewall ->
    ... -> IP NAT -> IP firewall -> { IP }  -> IP firewall -> IP NAT -> ...
     -> "layer2" IP firewall -> "layer2" IP NAT -> L2 firewall -> [OUTPUT]

That would be symmetric processing: we go through steps A through E on
input, and then E' through A' on output.  That's easy to understand.

I don't understand doing A-E on input, and then E', D', A', C', B' on
output, which is what you've documented.  Is there a reason for this?

(This is my only real sticking point right now.)

> > 3.  We're defining these bits of syntax ourselves, and we're expecting
> >     that administrators are going to rely on them for the security of
> >     their systems.  Given that, is "Volatile" the right classification
> >     for the new "family ether" and "layer2" configuration keywords?
> 
> We'll think more about this.

Please just make it "Committed."

> > 4.  Why is hpe_hdrinfo "void *" rather than "hook_pkt_info_t *"?  
> 
> For MAC layer, it is actually "mac_header_info_t *".

Oh ... ok; the documentation should probably make that clearer.

> > Void
> >     pointers are mildly evil, as they prevent the compiler and lint
> >     from doing their type-checking jobs.  What else would hpe_hdrinfo
> >     point to?
> 
> The major reason for using "void *" here, is that hook_pkt_event_t
> is a quite general structure that works with all protocols, thus I
> feel it is probably not appropriate to use types that specific to
> one protocol, e.g. mac_header_info_t is specific to MAC. Also it
> could point to other data structure for protocol other than MAC
> layer, though there's not such need right now.

I see.  I would have used a discriminated type here to make accidents
less likely (an enum + union, or just overlaid structures as is done
with sockaddr), but I guess this is your call.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Kais.Belgaied@sun.com Wed Dec 10 09:58:24 2008
Received: from sunmail3mpk.sfbay.sun.com (sunmail3mpk.SFBay.Sun.COM [129.146.11.52])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBAHwOWE016541
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 10 Dec 2008 09:58:24 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail3mpk.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mBAHwMtn005172;
	Wed, 10 Dec 2008 09:58:22 -0800 (PST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBO00F059X99800@brm-avmta-1.central.sun.com>; Wed,
 10 Dec 2008 10:58:21 -0700 (MST)
Received: from jurassic-x4600.sfbay.sun.com ([129.146.17.63])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBO009ZE9X8RB40@brm-avmta-1.central.sun.com>; Wed,
 10 Dec 2008 10:58:21 -0700 (MST)
Received: from [129.146.11.146]
 (sr1-jurassic-03.SFBay.Sun.COM [129.146.11.146])	by
 jurassic-x4600.sfbay.sun.com (8.14.3+Sun/8.14.3) with ESMTP id mBAHwK7e415669;
 Wed, 10 Dec 2008 09:58:20 -0800 (PST)
Date: Wed, 10 Dec 2008 09:58:20 -0800
From: Kais Belgaied <Kais.Belgaied@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <18751.60872.537103.651790@gargle.gargle.HOWL>
To: Darren Reed <Darren.Reed@sun.com>
Cc: Zhijun Fu <Zhijun.Fu@sun.com>, PSARC-ext@sun.com
Reply-to: Kais.Belgaied@sun.com
Message-id: <4940033C.10007@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; charset=ISO-8859-1; format=flowed
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
 <18751.55735.419739.64295@gargle.gargle.HOWL> <493FE9E3.9000005@sun.com>
 <18751.60872.537103.651790@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.16 (X11/20080807)
Status: RO
Content-Length: 337

There is a rather long discussion about the design and some 
architectural questions in parallel, outside this alias.
I was hoping that discussion converges before the case times out.

Darren, this case needs to be put back in waiting need spec, at least 
until the variety of
changes being proposes over the last week settle.

    Kais

From Zhijun.Fu@sun.com Thu Dec 11 03:26:51 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBBBQpMR011814
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 11 Dec 2008 03:26:51 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mBBBQenP016577
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Thu, 11 Dec 2008 04:26:51 -0700 (MST)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBP00D0HMGQTX00@nwk-avmta-1.sfbay.Sun.COM> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Thu, 11 Dec 2008 03:26:50 -0800 (PST)
Received: from sineb-mail-2.sun.com ([192.18.19.7])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBP004CSMGPLI40@nwk-avmta-1.sfbay.Sun.COM> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Thu,
 11 Dec 2008 03:26:50 -0800 (PST)
Received: from fe-apac-06.sun.com
 (fe-apac-06.sun.com [192.18.19.177] (may be forged))
	by sineb-mail-2.sun.com (8.13.6+Sun/8.12.9) with ESMTP id mBBBQnLr013036	for
 <PSARC-ext@sun.com>; Thu, 11 Dec 2008 11:26:49 +0000 (GMT)
Received: from conversion-daemon.mail-apac.sun.com by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 id <0KBP00201MAZQP00@mail-apac.sun.com> (original mail from Zhijun.Fu@Sun.COM)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Thu,
 11 Dec 2008 19:26:48 +0800 (SGT)
Received: from [129.158.215.37] by mail-apac.sun.com
 (Sun Java System Messaging Server 6.2-6.01 (built Apr  3 2006))
 with ESMTPSA id <0KBP004RAMGIHY10@mail-apac.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Thu, 11 Dec 2008 19:26:48 +0800 (SGT)
Date: Thu, 11 Dec 2008 19:22:28 +0800
From: Zhijun Fu <Zhijun.Fu@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <18751.60872.537103.651790@gargle.gargle.HOWL>
Sender: Zhijun.Fu@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: Darren Reed <Darren.Reed@sun.com>, PSARC-ext@sun.com
Reply-to: Zhijun.Fu@sun.com
Message-id: <4940F7F4.2030505@Sun.COM>
MIME-version: 1.0
Content-type: multipart/alternative;
 boundary="Boundary_(ID_odwJzwfi0gbIHoo/iDIaKg)"
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
 <18751.55735.419739.64295@gargle.gargle.HOWL> <493FE9E3.9000005@sun.com>
 <18751.60872.537103.651790@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.6 (X11/20071119)
Status: RO
Content-Length: 8707

This is a multi-part message in MIME format.

--Boundary_(ID_odwJzwfi0gbIHoo/iDIaKg)
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT

James Carlson wrote:
>>> 2.  You give the new rule processing as:
>>>       
> [...]
>   
>>>     I don't understand why the ordering of operations isn't just
>>>     reversed on output.  Assuming the input order is correct (and it
>>>     appears to me to be right), the "L2 firewall" element should be
>>>     the last operation before "[OUTPUT]."
>>>       
>> It could be a useful feature to combine filtering on L2 and L3
>> together, or even trigger conditional IP NAT from a L2 filtering rule,
>> though they are not supported in the current project scope, down the
>> road we may want to add it, and these features require that L2 firewall
>> to be processed before those "layer2" IP rules.
>>
>> Also, since the packet is received at MAC layer, it might be more
>> straight forward to do the parsing and matching from the MAC header,
>> as we need to parse the MAC header first anyway to locate the IP header
>> start. It also allows a simpler implementation in IPFilter but that
>> can be considered an implementation issue anyway.
>>
>> Would the above make sense?
>>     
>
> I'm not sure, but I don't think any of it answers the actual question
> I had.  You currently have this:
>
>      [INPUT] -> L2 firewall -> "layer2" IP NAT -> "layer2" IP firewall ->
>     ... -> IP NAT -> IP firewall -> { IP }  -> IP firewall -> IP NAT -> ...
>      -> L2 firewall -> "layer2" IP firewall -> "layer2" IP NAT -> [OUTPUT]
>
> But that's not entirely symmetric, and I don't see why.  I would have
> expected it to be:
>
>      [INPUT] -> L2 firewall -> "layer2" IP NAT -> "layer2" IP firewall ->
>     ... -> IP NAT -> IP firewall -> { IP }  -> IP firewall -> IP NAT -> ...
>      -> "layer2" IP firewall -> "layer2" IP NAT -> L2 firewall -> [OUTPUT]
>
> That would be symmetric processing: we go through steps A through E on
> input, and then E' through A' on output.  That's easy to understand.
>
> I don't understand doing A-E on input, and then E', D', A', C', B' on
> output, which is what you've documented.  Is there a reason for this?
>
> (This is my only real sticking point right now.

Jim,

Yes I know what you're saying. :-)

I agree that it makes sense for the processing to be entirely symmetric.
The problem with that, as I understand, is it will remove future chances
to combine L2 + L3 filtering together, or trigger L3 NAT conditionally
from L2 filtering rules.

(forgive me about the invented keywords below as it is just an example,
and won't be provided in this project)

For input, we might want to do something like below in the future:

#pass in family ether from 11:22:33:44:55:55 to any l2-head 100
#pass in proto tcp from any to any l2-group 100 layer2

For output, we might want to do something similar:

#pass out family ether from 66:55:44:33:22:11 to any l2-head 200
#pass out proto tcp from any to any l2-group 200 layer2

(the same for combining L2 filtering + "layer2" L3 NAT)

To do this, we need the L2 firewall to be processed earlier
than "layer2" IP firewall (and "layer2" IP NAT), for
both INPUT and OUTPUT, as otherwise we don't know whether
a "layer2" IP firewall/NAT rule should be processed or not
if we do E -> D -> C -> B -> A for output, as the rules
can be conditional which depend on the L2 firewall rules,
which haven't be processed at that time.

Is there actual problems in your mind that could be caused by
the current "not entirely symmetric" approach?


Thank you,

Zhijun

-- 
#mdb -K
[0]> eri.prc.sun.com::walk staff s|::print staff_t s_name|
::grep .== zhijun |::eval <s=K|::print staff_t
Zhijun.Fu@Sun.COM, x84349
Network Virtualization & Performance Team,
Solaris Core Operating Systems
Since Jul 10,2006
[0]> :c


--Boundary_(ID_odwJzwfi0gbIHoo/iDIaKg)
Content-type: text/html; charset=ISO-8859-1
Content-transfer-encoding: 7BIT

<!DOCTYPE html PUBLIC "-//W3C//DTD HTML 4.01 Transitional//EN">
<html>
<head>
  <meta content="text/html;charset=ISO-8859-1" http-equiv="Content-Type">
  <title></title>
</head>
<body bgcolor="#ffffff" text="#000000">
James Carlson wrote:<br>
<blockquote cite="mid:18751.60872.537103.651790@gargle.gargle.HOWL"
 type="cite">
  <blockquote type="cite">
    <blockquote type="cite">
      <pre wrap="">2.  You give the new rule processing as:
      </pre>
    </blockquote>
  </blockquote>
  <pre wrap=""><!---->[...]
  </pre>
  <blockquote type="cite">
    <blockquote type="cite">
      <pre wrap="">    I don't understand why the ordering of operations isn't just
    reversed on output.  Assuming the input order is correct (and it
    appears to me to be right), the "L2 firewall" element should be
    the last operation before "[OUTPUT]."
      </pre>
    </blockquote>
    <pre wrap="">It could be a useful feature to combine filtering on L2 and L3
together, or even trigger conditional IP NAT from a L2 filtering rule,
though they are not supported in the current project scope, down the
road we may want to add it, and these features require that L2 firewall
to be processed before those "layer2" IP rules.

Also, since the packet is received at MAC layer, it might be more
straight forward to do the parsing and matching from the MAC header,
as we need to parse the MAC header first anyway to locate the IP header
start. It also allows a simpler implementation in IPFilter but that
can be considered an implementation issue anyway.

Would the above make sense?
    </pre>
  </blockquote>
  <pre wrap=""><!---->
I'm not sure, but I don't think any of it answers the actual question
I had.  You currently have this:

     [INPUT] -&gt; L2 firewall -&gt; "layer2" IP NAT -&gt; "layer2" IP firewall -&gt;
    ... -&gt; IP NAT -&gt; IP firewall -&gt; { IP }  -&gt; IP firewall -&gt; IP NAT -&gt; ...
     -&gt; L2 firewall -&gt; "layer2" IP firewall -&gt; "layer2" IP NAT -&gt; [OUTPUT]

But that's not entirely symmetric, and I don't see why.  I would have
expected it to be:

     [INPUT] -&gt; L2 firewall -&gt; "layer2" IP NAT -&gt; "layer2" IP firewall -&gt;
    ... -&gt; IP NAT -&gt; IP firewall -&gt; { IP }  -&gt; IP firewall -&gt; IP NAT -&gt; ...
     -&gt; "layer2" IP firewall -&gt; "layer2" IP NAT -&gt; L2 firewall -&gt; [OUTPUT]

That would be symmetric processing: we go through steps A through E on
input, and then E' through A' on output.  That's easy to understand.

I don't understand doing A-E on input, and then E', D', A', C', B' on
output, which is what you've documented.  Is there a reason for this?

(This is my only real sticking point right now.</pre>
</blockquote>
<tt><br>
Jim,<br>
<br>
Yes I know what you're saying.<span class="moz-smiley-s1"><span> :-) </span></span><br>
<br>
I agree that it makes sense for the processing to be entirely symmetric.<br>
The problem with that, as I understand, is it will remove future chances<br>
to combine L2 + L3 filtering together, or trigger L3 NAT conditionally<br>
from L2 filtering rules.<br>
<br>
</tt><tt>(forgive me about the invented keywords below as it is just an
example,<br>
and won't be provided in this project)<br>
<br>
</tt><tt>For input, we might want to do something like below in the
future:<br>
<br>
#pass in family ether from 11:22:33:44:55:55 to any l2-head 100<br>
#pass in proto tcp from any to any l2-group 100 layer2<br>
<br>
For output, we might want to do something similar:<br>
</tt><tt><br>
#pass out family ether from 66:55:44:33:22:11 to any l2-head 200<br>
#pass out proto tcp from any to any l2-group 200 layer2</tt><br>
<tt><br>
(the same for combining L2 filtering + "layer2" L3 NAT)<br>
<br>
To do this, we need the L2 firewall to be processed earlier<br>
than "layer2" IP firewall (and "layer2" IP NAT), for<br>
both INPUT and OUTPUT, as otherwise we don't know whether<br>
a "layer2" IP firewall/NAT rule should be processed or not<br>
if we do E -&gt; D -&gt; C -&gt; B -&gt; A for output, as the rules<br>
can be conditional which depend on the L2 firewall rules,<br>
which haven't be processed at that time.<br>
<br>
Is there actual problems in your mind that could be caused by<br>
the current "not entirely symmetric" approach?<br>
<br>
<br>
Thank you,<br>
<br>
Zhijun<br>
</tt><br>
<pre class="moz-signature" cols="72">-- 
#mdb -K
[0]&gt; eri.prc.sun.com::walk staff s|::print staff_t s_name|
::grep .== zhijun |::eval &lt;s=K|::print staff_t
<a class="moz-txt-link-abbreviated" href="mailto:Zhijun.Fu@Sun.COM">Zhijun.Fu@Sun.COM</a>, x84349
Network Virtualization &amp; Performance Team,
Solaris Core Operating Systems
Since Jul 10,2006
[0]&gt; :c
</pre>
</body>
</html>

--Boundary_(ID_odwJzwfi0gbIHoo/iDIaKg)--

From erik.nordmark@sun.com Thu Dec 11 11:06:49 2008
Received: from sunmail3mpk.sfbay.sun.com (sunmail3mpk.SFBay.Sun.COM [129.146.11.52])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBBJ6nf7023846
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 11 Dec 2008 11:06:49 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail3mpk.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mBBJ6hGr011034;
	Thu, 11 Dec 2008 11:06:44 -0800 (PST)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBQ001337R67500@nwk-avmta-1.sfbay.Sun.COM>; Thu,
 11 Dec 2008 11:06:42 -0800 (PST)
Received: from jurassic.eng.sun.com ([129.146.224.31])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBQ00K177R3WV60@nwk-avmta-1.sfbay.Sun.COM>; Thu,
 11 Dec 2008 11:06:39 -0800 (PST)
Received: from [10.7.251.248] (punchin-nordmark.SFBay.Sun.COM [10.7.251.248])
	by jurassic.eng.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBBJ6Ze2949081
	(version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-SHA bits=256 verify=NO); Thu,
 11 Dec 2008 11:06:36 -0800 (PST)
Date: Thu, 11 Dec 2008 11:06:35 -0800
From: Erik Nordmark <erik.nordmark@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <4940F7F4.2030505@Sun.COM>
To: Zhijun.Fu@sun.com
Cc: James Carlson <James.D.Carlson@sun.com>, Darren Reed <Darren.Reed@sun.com>,
        PSARC-ext@sun.com
Message-id: <494164BB.9050608@sun.com>
MIME-version: 1.0
Content-type: text/plain; charset=ISO-8859-1; format=flowed
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
 <18751.55735.419739.64295@gargle.gargle.HOWL> <493FE9E3.9000005@sun.com>
 <18751.60872.537103.651790@gargle.gargle.HOWL> <4940F7F4.2030505@Sun.COM>
User-Agent: Thunderbird 2.0.0.17 (X11/20081023)
Status: RO
Content-Length: 1381

Zhijun Fu wrote:

> For input, we might want to do something like below in the future:
> 
> #pass in family ether from 11:22:33:44:55:55 to any l2-head 100
> #pass in proto tcp from any to any l2-group 100 layer2
> 
> For output, we might want to do something similar:
> 
> #pass out family ether from 66:55:44:33:22:11 to any l2-head 200
> #pass out proto tcp from any to any l2-group 200 layer2
> 
> (the same for combining L2 filtering + "layer2" L3 NAT)
> 
> To do this, we need the L2 firewall to be processed earlier
> than "layer2" IP firewall (and "layer2" IP NAT), for
> both INPUT and OUTPUT, as otherwise we don't know whether
> a "layer2" IP firewall/NAT rule should be processed or not
> if we do E -> D -> C -> B -> A for output, as the rules
> can be conditional which depend on the L2 firewall rules,
> which haven't be processed at that time.

One way to think about this is that what you have as the rule with 
l2-head isn't a traditional firewall rule, but that it instead is a 
classification rule whose result is to tag/label the packet for further 
processing.
Other rules can then be written which use the tag/label.

In that case clearly the classification has to happen before its use.

But I don't see l2-head and l2-group in the current case, thus to avoid 
confusing the users couldn't we simplify the current description to be 
symmetric?

    Erik




From carlsonj@phorcys.east.sun.com Thu Dec 11 11:45:09 2008
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBBJj9Lm024444
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 11 Dec 2008 11:45:09 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mBBJj2vW056955;
	Thu, 11 Dec 2008 12:45:05 -0700 (MST)
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBQ0062D9J3TZ00@brm-avmta-1.central.sun.com>; Thu,
 11 Dec 2008 12:45:03 -0700 (MST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBQ00KX59J23QA0@brm-avmta-1.central.sun.com>; Thu,
 11 Dec 2008 12:45:02 -0700 (MST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id mBBJj22Q024486; Thu,
 11 Dec 2008 14:45:02 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id mBBJj2Lm024483; Thu,
 11 Dec 2008 14:45:02 -0500 (EST)
Date: Thu, 11 Dec 2008 14:45:02 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <494164BB.9050608@sun.com>
To: Erik Nordmark <Erik.Nordmark@sun.com>
Cc: Zhijun.Fu@sun.com, PSARC-ext@sun.com, Darren Reed <Darren.Reed@sun.com>
Message-id: <18753.28094.430059.624736@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
 <18751.55735.419739.64295@gargle.gargle.HOWL> <493FE9E3.9000005@sun.com>
 <18751.60872.537103.651790@gargle.gargle.HOWL> <4940F7F4.2030505@Sun.COM>
 <494164BB.9050608@sun.com>
Status: RO
Content-Length: 1949

Erik Nordmark writes:
> Zhijun Fu wrote:
> > To do this, we need the L2 firewall to be processed earlier
> > than "layer2" IP firewall (and "layer2" IP NAT), for
> > both INPUT and OUTPUT, as otherwise we don't know whether
> > a "layer2" IP firewall/NAT rule should be processed or not
> > if we do E -> D -> C -> B -> A for output, as the rules
> > can be conditional which depend on the L2 firewall rules,
> > which haven't be processed at that time.
> 
> One way to think about this is that what you have as the rule with 
> l2-head isn't a traditional firewall rule, but that it instead is a 
> classification rule whose result is to tag/label the packet for further 
> processing.
> Other rules can then be written which use the tag/label.

Yes.  And I think there's probably a more flexible way to do this
using the existing "head" logic in IP filter, and just making the
internal logic smarter when dealing with packets that may have either
L2 or L3 origin.

> In that case clearly the classification has to happen before its use.
> 
> But I don't see l2-head and l2-group in the current case, thus to avoid 
> confusing the users couldn't we simplify the current description to be 
> symmetric?

I think the reason the submitter wants to do this is to retain the
option of doing the same "l2-head" thing as was originally proposed,
just at some later date.

As for the current proposal, I think symmetric is much easier to
understand.  It would be really strange to find that (for instance) an
L2 firewall rule precluding sending packets to a particular MAC
destination could be overridden by an L2 IP firewall or IP NAT rule
that rewrites the IP destination, but that the input side blocks the
traffic as expected.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Darren.Reed@sun.com Thu Dec 11 18:17:16 2008
Received: from sunmail3mpk.sfbay.sun.com (sunmail3mpk.SFBay.Sun.COM [129.146.11.52])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBC2HGik024363
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 11 Dec 2008 18:17:16 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail3mpk.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id mBC2HD10025983
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Thu, 11 Dec 2008 18:17:14 -0800 (PST)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBQ00A0JROQT400@nwk-avmta-1.sfbay.Sun.COM> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Thu, 11 Dec 2008 18:17:14 -0800 (PST)
Received: from gmp-eb-inf-1.sun.com ([192.18.6.21])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBQ00CC8ROP1CE0@nwk-avmta-1.sfbay.Sun.COM> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Thu,
 11 Dec 2008 18:17:13 -0800 (PST)
Received: from fe-emea-09.sun.com (gmp-eb-lb-2-fe2.eu.sun.com [192.18.6.11])
	by gmp-eb-inf-1.sun.com (8.13.7+Sun/8.12.9) with ESMTP id mBC2HC5C024851	for
 <PSARC-ext@sun.com>; Fri, 12 Dec 2008 02:17:12 +0000 (GMT)
Received: from conversion-daemon.fe-emea-09.sun.com by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KBQ00501RG3A000@fe-emea-09.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Fri,
 12 Dec 2008 02:17:12 +0000 (GMT)
Received: from [129.158.90.33] by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KBQ000UWROMPIA0@fe-emea-09.sun.com>; Fri,
 12 Dec 2008 02:17:12 +0000 (GMT)
Date: Fri, 12 Dec 2008 13:17:04 +1100
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <4940033C.10007@Sun.COM>
Sender: Darren.Reed@sun.com
To: Kais.Belgaied@sun.com
Cc: Zhijun Fu <Zhijun.Fu@sun.com>, PSARC-ext@sun.com
Message-id: <4941C9A0.70903@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
 <18751.55735.419739.64295@gargle.gargle.HOWL> <493FE9E3.9000005@sun.com>
 <18751.60872.537103.651790@gargle.gargle.HOWL> <4940033C.10007@Sun.COM>
User-Agent: Thunderbird 2.0.0.18 (Windows/20081105)
Status: RO
Content-Length: 425

Kais Belgaied wrote:
> There is a rather long discussion about the design and some 
> architectural questions in parallel, outside this alias.
> I was hoping that discussion converges before the case times out.
>
> Darren, this case needs to be put back in waiting need spec, at least 
> until the variety of
> changes being proposes over the last week settle.

Will do.  Sorry for not picking this email up sooner.

Darren


From Darren.Reed@sun.com Thu Dec 11 18:33:15 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBC2XEIE024450
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 11 Dec 2008 18:33:14 -0800 (PST)
Received: from brm-avmta-1.central.sun.com (brm-avmta-1.Central.Sun.COM [129.147.4.11])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id mBC2XAHh018987
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Fri, 12 Dec 2008 02:33:13 GMT
Received: from pmxchannel-daemon.brm-avmta-1.central.sun.com by
 brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBQ00103SFBA700@brm-avmta-1.central.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Thu, 11 Dec 2008 19:33:11 -0700 (MST)
Received: from gmp-eb-inf-1.sun.com ([192.18.6.21])
 by brm-avmta-1.central.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBQ007ZLSFA4VB0@brm-avmta-1.central.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Thu,
 11 Dec 2008 19:33:10 -0700 (MST)
Received: from fe-emea-10.sun.com (gmp-eb-lb-2-fe2.eu.sun.com [192.18.6.11])
	by gmp-eb-inf-1.sun.com (8.13.7+Sun/8.12.9) with ESMTP id mBC2XAOp025188	for
 <PSARC-ext@sun.com>; Fri, 12 Dec 2008 02:33:10 +0000 (GMT)
Received: from conversion-daemon.fe-emea-10.sun.com by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KBQ00C01SELMQ00@fe-emea-10.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Fri,
 12 Dec 2008 02:33:09 +0000 (GMT)
Received: from [129.158.90.33] by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KBQ00E1YSF60N80@fe-emea-10.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Fri, 12 Dec 2008 02:33:09 +0000 (GMT)
Date: Fri, 12 Dec 2008 13:33:01 +1100
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <18751.60872.537103.651790@gargle.gargle.HOWL>
Sender: Darren.Reed@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: Zhijun Fu <Zhijun.Fu@sun.com>, PSARC-ext@sun.com
Message-id: <4941CD5D.2060106@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
 <18751.55735.419739.64295@gargle.gargle.HOWL> <493FE9E3.9000005@sun.com>
 <18751.60872.537103.651790@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.18 (Windows/20081105)
Status: RO
Content-Length: 1204

James Carlson wrote:
> ...
>>> 3.  We're defining these bits of syntax ourselves, and we're expecting
>>>     that administrators are going to rely on them for the security of
>>>     their systems.  Given that, is "Volatile" the right classification
>>>     for the new "family ether" and "layer2" configuration keywords?
>>>       
>> We'll think more about this.
>>     
>
> Please just make it "Committed."
>   

For "family ether", I've no problem with "Committed."

The "layer2" bits I consider to be a blight on the configruation syntax,
not to mention that implementation atrocity that results in policy needing
to be defined twice, and I will be looking for a way to arcitect it out in
the future.

Whilst it might appeal to you (since you pretty much got in the way of
anything else), it really does not fit into anything futurish for 
ipfilter. It is
a dead end piece of syntax and we should not be carrying that baggage
around for any longer than it needs to be.

Having to put up with:
* layer 2 rules
* layer 3 rules
* layer 3 rules for layer 2
is something that we can accomdate in the short term for the sake of
expediency but in long term, the last of those three needs to die.

Darren


From carlsonj@phorcys.east.sun.com Fri Dec 12 04:38:39 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBCCcdXt021468
	for <psarc-ext@sac.sfbay.sun.com>; Fri, 12 Dec 2008 04:38:39 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id mBCCcOw7023013;
	Fri, 12 Dec 2008 12:38:35 GMT
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBR00305KGAKO00@nwk-avmta-1.sfbay.Sun.COM>; Fri,
 12 Dec 2008 04:38:34 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBR00BRYKG996D0@nwk-avmta-1.sfbay.Sun.COM>; Fri,
 12 Dec 2008 04:38:33 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id mBCCcXmZ026149; Fri,
 12 Dec 2008 07:38:33 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id mBCCcXNS026146; Fri,
 12 Dec 2008 07:38:33 -0500 (EST)
Date: Fri, 12 Dec 2008 07:38:33 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <4941CD5D.2060106@Sun.COM>
To: Darren Reed <Darren.Reed@sun.com>
Cc: Zhijun Fu <Zhijun.Fu@sun.com>, PSARC-ext@sun.com
Message-id: <18754.23369.185448.834248@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
 <18751.55735.419739.64295@gargle.gargle.HOWL> <493FE9E3.9000005@sun.com>
 <18751.60872.537103.651790@gargle.gargle.HOWL> <4941CD5D.2060106@Sun.COM>
Status: RO
Content-Length: 1613

Darren Reed writes:
> The "layer2" bits I consider to be a blight on the configruation syntax,
> not to mention that implementation atrocity that results in policy needing
> to be defined twice, and I will be looking for a way to arcitect it out in
> the future.

The problem I'm pointing out here is that it is incongruous to make
crucial security configuration syntax "Volatile."  If there's anything
I don't want to have disappear or change in meaning over time, it'd
have to be my system security configuration.

I agree that having to specify "I want this L3 rule to run at L2" or
more generally having go specify which hook to use for a given rule
seems quite wrong.

> Whilst it might appeal to you (since you pretty much got in the way of
> anything else), it really does not fit into anything futurish for 
> ipfilter.

So much for civility.

> is something that we can accomdate in the short term for the sake of
> expediency but in long term, the last of those three needs to die.

Then this case is incomplete.  It needs to explain where we're going
and how we'll get there.

Do we need to distinguish between the "layer 3" and "layer 3 at L2"
cases, and, if we do, how do we do that in a way that will not just
result in future breakage?

If you have to patch it on for now, that's ok, but please do explain
how we get from the patched-on state to a longer-term usable state.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Darren.Reed@sun.com Sun Dec 14 18:30:51 2008
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBF2UocY009779
	for <psarc-ext@sac.sfbay.Sun.COM>; Sun, 14 Dec 2008 18:30:51 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id mBF2UfG8010282
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Mon, 15 Dec 2008 10:30:49 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBW00C01CB95Y00@nwk-avmta-2.sfbay.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Sun, 14 Dec 2008 18:30:45 -0800 (PST)
Received: from gmp-eb-inf-1.sun.com ([192.18.6.21])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBW00ABTCB8YUD0@nwk-avmta-2.sfbay.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Sun,
 14 Dec 2008 18:30:45 -0800 (PST)
Received: from fe-emea-10.sun.com (gmp-eb-lb-1-fe3.eu.sun.com [192.18.6.10])
	by gmp-eb-inf-1.sun.com (8.13.7+Sun/8.12.9) with ESMTP id mBF2UhBj009179	for
 <PSARC-ext@sun.com>; Mon, 15 Dec 2008 02:30:43 +0000 (GMT)
Received: from conversion-daemon.fe-emea-10.sun.com by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KBW00D01C8QLY00@fe-emea-10.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Mon,
 15 Dec 2008 02:30:43 +0000 (GMT)
Received: from [129.158.90.33] by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KBW00I4RCB59980@fe-emea-10.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Mon, 15 Dec 2008 02:30:43 +0000 (GMT)
Date: Mon, 15 Dec 2008 13:30:39 +1100
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <18754.23369.185448.834248@gargle.gargle.HOWL>
Sender: Darren.Reed@sun.com
To: James Carlson <James.D.Carlson@sun.com>
Cc: Zhijun Fu <Zhijun.Fu@sun.com>, PSARC-ext@sun.com
Message-id: <4945C14F.7090409@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
 <18751.55735.419739.64295@gargle.gargle.HOWL> <493FE9E3.9000005@sun.com>
 <18751.60872.537103.651790@gargle.gargle.HOWL> <4941CD5D.2060106@Sun.COM>
 <18754.23369.185448.834248@gargle.gargle.HOWL>
User-Agent: Thunderbird 2.0.0.18 (Windows/20081105)
Status: RO
Content-Length: 1759

James Carlson wrote:
> Darren Reed writes:
>   
>> The "layer2" bits I consider to be a blight on the configruation syntax,
>> not to mention that implementation atrocity that results in policy needing
>> to be defined twice, and I will be looking for a way to arcitect it out in
>> the future.
>>     
>
> The problem I'm pointing out here is that it is incongruous to make
> crucial security configuration syntax "Volatile."  If there's anything
> I don't want to have disappear or change in meaning over time, it'd
> have to be my system security configuration.
>   

...

>> is something that we can accomdate in the short term for the sake of
>> expediency but in long term, the last of those three needs to die.
>>     
>
> Then this case is incomplete.  It needs to explain where we're going
> and how we'll get there.
>
> Do we need to distinguish between the "layer 3" and "layer 3 at L2"
> cases, and, if we do, how do we do that in a way that will not just
> result in future breakage?
>
> If you have to patch it on for now, that's ok, but please do explain
> how we get from the patched-on state to a longer-term usable state.
>   

In the fullness of time, IPFilter will allow administrators to
"decapsulate" packets, so that in instances where there are
interesting IP headers "inside" the packet in clear text, it
will be possible to filter on those.

Thus filtering on layer 3 headers from layer 2 should just
become another use of that design rather than something
special.

The "layer2" tag was adopted after you expressed distate for
having an "ip-head' or "l2-head" option with the ipf rules.
Given that the "layer2" keyword does not fit at all with future
direction, the only option is to flag it as "volatile" (or "obsolete.")

Darren


From carlsonj@phorcys.east.sun.com Mon Dec 15 04:27:32 2008
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id mBFCRWZi006780
	for <psarc-ext@sac.sfbay.sun.com>; Mon, 15 Dec 2008 04:27:32 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id mBFCR4FI016658;
	Mon, 15 Dec 2008 12:27:28 GMT
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KBX005053XLYJ00@nwk-avmta-2.sfbay.sun.com>; Mon,
 15 Dec 2008 04:27:21 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KBX001TX3XKR7B0@nwk-avmta-2.sfbay.sun.com>; Mon,
 15 Dec 2008 04:27:21 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id mBFCRK6g001294; Mon,
 15 Dec 2008 07:27:20 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id mBFCRKXK001291; Mon,
 15 Dec 2008 07:27:20 -0500 (EST)
Date: Mon, 15 Dec 2008 07:27:20 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: PSARC/2008/249 - Packet Interception for the MAC layer
In-reply-to: <4945C14F.7090409@Sun.COM>
To: Darren Reed <Darren.Reed@sun.com>
Cc: Zhijun Fu <Zhijun.Fu@sun.com>, PSARC-ext@sun.com
Message-id: <18758.19752.656115.799036@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <4927E0F8.8080103@Sun.COM> <493765A9.7080103@Sun.COM>
 <18751.55735.419739.64295@gargle.gargle.HOWL> <493FE9E3.9000005@sun.com>
 <18751.60872.537103.651790@gargle.gargle.HOWL> <4941CD5D.2060106@Sun.COM>
 <18754.23369.185448.834248@gargle.gargle.HOWL> <4945C14F.7090409@Sun.COM>
Status: RO
Content-Length: 1671

Darren Reed writes:
> > If you have to patch it on for now, that's ok, but please do explain
> > how we get from the patched-on state to a longer-term usable state.
> >   
> 
> In the fullness of time, IPFilter will allow administrators to
> "decapsulate" packets, so that in instances where there are
> interesting IP headers "inside" the packet in clear text, it
> will be possible to filter on those.
> 
> Thus filtering on layer 3 headers from layer 2 should just
> become another use of that design rather than something
> special.

So, that future design won't use the "layer2" tag, correct?

> The "layer2" tag was adopted after you expressed distate for
> having an "ip-head' or "l2-head" option with the ipf rules.
> Given that the "layer2" keyword does not fit at all with future
> direction, the only option is to flag it as "volatile" (or "obsolete.")

I expressed distaste for "ip-head" because it forces the user to work
around design issues in IP Filter itself, telling the system when to
advance the pointer from the MAC layer to the network layer header,
rather than just using 'head' for grouping.  (And I'm not sure what
you mean, because the original specification from last April has a
"layer2" tag, and I don't _think_ we talked about it before that.)

Given that "layer2" is a temporary scheme, it sounds like "Obsolete
Volatile" is the right way to go.  The man page should warn that the
keyword may go away in the future.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Darren.Reed@sun.com Tue Jan 20 07:32:56 2009
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n0KFWteq028403
	for <psarc-ext@sac.sfbay.sun.com>; Tue, 20 Jan 2009 07:32:55 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id n0KFWabQ006905
	for <@sunmail2sca.sfbay.sun.com:psarc-ext@sun.com>; Tue, 20 Jan 2009 15:32:54 GMT
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KDS00F590ITYB00@nwk-avmta-1.sfbay.Sun.COM> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Tue, 20 Jan 2009 07:32:53 -0800 (PST)
Received: from gmp-eb-inf-1.sun.com ([192.18.6.21])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KDS00DU10IQPOC0@nwk-avmta-1.sfbay.Sun.COM> for
 psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Tue,
 20 Jan 2009 07:32:51 -0800 (PST)
Received: from fe-emea-09.sun.com (gmp-eb-lb-1-fe3.eu.sun.com [192.18.6.10])
	by gmp-eb-inf-1.sun.com (8.13.7+Sun/8.12.9) with ESMTP id n0KFWoge009054	for
 <psarc-ext@sun.com>; Tue, 20 Jan 2009 15:32:50 +0000 (GMT)
Received: from conversion-daemon.fe-emea-09.sun.com by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KDS0040101OCF00@fe-emea-09.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for psarc-ext@sun.com (ORCPT psarc-ext@sun.com); Tue,
 20 Jan 2009 15:32:50 +0000 (GMT)
Received: from [129.157.18.18] by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KDS00JDI0IKJRE0@fe-emea-09.sun.com> for psarc-ext@sun.com
 (ORCPT psarc-ext@sun.com); Tue, 20 Jan 2009 15:32:44 +0000 (GMT)
Date: Tue, 20 Jan 2009 16:31:05 +0100
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <47FC5A55.2050405@Sun.COM>
Sender: Darren.Reed@sun.com
To: PSARC-EXT <PSARC-ext@sun.com>
Cc: Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <4975EE39.9080906@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM>
User-Agent: Thunderbird 2.0.0.19 (Windows/20081209)
Status: RO
Content-Length: 23094

I'm restarting the timer for this case.

The updated diff-marked specification is attached below.

Darren
--------------------------------------------------------------

Abstract
========
This case will extend PSARC/2005/334, by adding the ability to intercept
packets in MAC layer using the PFHooks infrastructure.

| This case only makes a few changes, an addition, to the interfaces that were
| committed to by PSARC/2008/219 (see "net_getlifaddr", "hook_pkt_event_t",
| and "new NIC event" below for details)

Release Biding
--------------
This case seeks for a patch binding.

Introduction
============
The PFHooks project, PSARC/2005/334, provide the ability to intercept packets
in IP layer by adding hooks into the network stack. 

Since its integration, there has been customer requirements for the ability
to intercept packets in MAC layer as well, also it is needed to enforce
security rules for xVM guest domains and exclusive zones.

Goals
-----
This case seeks to meet the following goals:
* provide the hooks in MAC layer that allows consumers to register on to 
  intercept packets;

* provide the netinfo interface for MAC layer that gives consumers access to
  interface information, and the ability to inject or emit packets directly;

* modify IPFilter to allow the administrator to specify layer 2 rules, which
  includes ethernet filtering rules and IP Filtering/NAT rules.

Out of scope
------------
This project only provides the ability to specify ethernet filtering rules
to match ethernet packets, and IP filtering/NAT rules to match/modify IP
packets at MAC layer. Providing the ability to specify rules to filter 
non-ethernet packets by matching the MAC header is out of scope for this
project.

The detailed design are described below for each major components.

Netinfo interface for MAC layer
===============================
netinfo interfaces
------------------
The hooks provided will generate events for NH_PHYSICAL_IN and NH_PHYSICAL_OUT,
using the same interface as IPv4 and IPv6 do in PSARC/2005/334.

The following functions will be supported through the netinfo(9f) framework:
net_getifname()
net_phylookup()
net_phygetnext()
net_getlifaddr()
net_inject()
net_getmtu()

All of the other functions in the netinfo(9f) framework will return a value
indicating that they are unsupported. The return values for the above 
functions only have meaning with the scope of the corresponding family - 
it is not correct to use a value returned by net_getifname() using the
ethernet net_data_t handle with net_phylookup() for IPv4.

The callback for these events will receive a pointer to a hook_pkt_event_t
structure that has the following fields filled out:

hpe_ifp - 0 for NH_PHYSICAL_OUT, otherwise a value indicating which
          interface the NH_PHYSICAL_IN event is associated with;
hpe_ofp - 0 for NH_PHYSICAL_IN, otherwise a value indicating which
          interface the NH_PHYSICAL_OUT event is associated with;
hpe_hdr - points to the start of the MAC header
hpe_mb  - points to the start of the mblk_t that holds hpe_hdr;
hpe_mp  - points to the mblk_t that is the start of the packet.
hpe_hpeinfo - points to mac_header_info_t which contains MAC header information

| net_getlifaddr
| --------------
| The net_getlifaddr() function returns the address for a given interface.
| For existing IP netinfo it returns the IP address, and for MAC layer netinfo
| it returns MAC address for an interface, the the usage of this function
| is slightly different in the two situations.
| 
| int net_getlifaddr(const net_data_t net, const phy_if_t ifp,
|     const net_if_t lif,  int const type, struct sockaddr *storage);
| 
|     net
|          value returned from a successful call to net_protocol_lookup.
| 
|     ifp
|          value returned from a successful call to net_phylookup
|          or net_phygetnext, indicating which network interface
|          the information should be returned from.
|
|     lif
|          indicating which logical interface to fetch the address from.
| 
|     type
|          this indicates what type of address should be returned.
| 
|     storage
|          pointer to an area of memory to store the address data.
| 
| This case introduces a slightly different usage for this function
| when used to retrieve MAC layer information. Unlike IP, MAC doesn't
| have the concept of logical interface, so the caller should pass in
| the physical interface as ifp, and pass in a 0 as the lif because
| there is no valid lif for MAC.
| 
| Each call to net_getlifaddr requires that the caller pass in
| a pointer to an array of address information types to retrieve
| and an accompanying pointer to an array of pointers to struct
| sockaddr_dl structures in which to copy the address information
| into. See below for an example of how to use this function.
| 
| Each member of the address type array should be one of NA_ADDRESS,
| NA_PEER, NA_BROADCAST or NA_NETMASK, and it is up to each layer 2
| protocol to implement the address type. For Ethernet, NA_ADDRESS
| and NA_BROADCAST are supported, and NA_BROADCAST always return
| ff:ff:ff:ff:ff:ff.
| 
hook_pkt_event_t
----------------
In order to intercept IP packets at MAC layer, IPFilter needs to know the 
size of the MAC header to locate the IP header start. The problems is the wifi
header is not self explained, parsing it requires information from mac handle
thus IPFilter cannot do the parsing itself, so we need to rely on the MAC
plugin to parse the header, pass the information through Hook framework to
IPFilter, so it can identify IP header start correctly.

While adding a header length field to hook_pkt_info_t solves the problem above,
down the road we may want to provide the ability to match wifi header, which
requires information of the wifi header fields in IPFilter, not just the header
| length, thus we propose to add a pointer to hook_pkt_event_t, which points at
a structure of mac_header_info_t, and pass this through the Hook framework,
so the hook consumers, like IPFilter, can have the needed information for
| the MAC header. The new hook_pkt_event_t strucuture would look like:

typedef struct hook_pkt_event {
        net_handle_t            hpe_protocol;
        phy_if_t                hpe_ifp;
        phy_if_t                hpe_ofp;
        void                    *hpe_hdr;
        mblk_t                  **hpe_mp;
        mblk_t                  *hpe_mb;
        int                     hpe_flags;
-       void                    *hpe_reserved[2];
+       void                    *hpe_hdrinfo;
+       void                    *hpe_reserved[1];
} hook_pkt_event_t;

For existing IP/ARP Hooks, the header format is self explained, so hpe_hdrinfo
will be NULL and IPFilter does the header parsing itself as before.

| MAC client index
| ----------------
| L2 filtering is based on MAC client which is introduced by Crossbow project,
| and the filtering is done on a per MAC client basis. When users specify a
| link name "net0", this corresponds to the traffic going through the primary
| MAC client of net0, e.g. IP on top of that data link. 
| 
| The MAC client index is introduced in this project, which uniquely identifies
| a MAC client and is used by the layer 2 netinfo interface in the same way
| as the ifindex is used by the IP netinfo interface. And layer 2 netinfo
| provides the mapping between data link name and index of the primary MAC
| client of that data link, through net_getifname() and net_phylookup().

new NIC event
-------------
The status of network in the operating system often changes, from unplugging
a system from network temporarily, to an interface's IP address changing
as a result of DHCP. Thus PFHooks framework provides event notification
mechanism for this.

The callback for these events will receive a pointer to a hook_nic_event_t
structure that has the following fields filled out:

hne_protocol - network protocol for events, returned from net_lookup

hne_nic      - physical interface associated with event

hne_lif      - logical interface (if any) associated with event

hne_event    - type of event occuring. The current list of events available is:

	NE_PLUMB
	       indicates that an interface has just been created

	NE_UNPLUMB
	       indicates that an interface has just been destroyed and that
	       no more events should be received for it

	NE_UP
	       indicates that an interface has changed state to "up" and 
               may now generate packet events.

	NE_DOWN
	       indicates that an interface has changed state to "down" and
	       will no longer generate packet events.

	NE_ADDRESS_CHANGE
	       indicates that an address on an interface has changed.

hne_data     - pointer to extra data about event or NULL if none

hne_datalen  - size of data pointed to by hne_data (can be 0)

NE_NAME_CHANGE event
~~~~~~~~~~~~~~~~~~~~
As Clearview UV (PSARC/2006/499, PSARC/2007/527, PSARC/2008/002) introduces
the ability to rename a data link, we need to capture this event in order to
update IPFilter rules accrodingly. Thus we propose an extension to
PSARC/2008/219 by adding a new hook event NE_NAME_CHANGE to nic_event_t
to indicate the that an interface has been renamed, and this particular event
is only available to layer 2 netinfo. In IP, changing of an interface name
is represented by a NE_UNPLUMB and NE_PLUMB event pair.

typedef enum nic_event {
         NE_PLUMB = 1,
         NE_UNPLUMB,
         NE_UP,
         NE_DOWN,
         NE_ADDRESS_CHANGE,
+        NE_NAME_CHANGE
} nic_event_t;

Design considerations
~~~~~~~~~~~~~~~~~~~~~
IPFilter rules always match by name, and only the current link names are used
for matching, not old names. Uppon NE_NAME_CHANGE event, IPFilter will walk
all the layer 2 rules, and resolve the interface name stored in the rule
structure into interface pointers. So when the link is renamed, rules using
old link names are invalidated, and rules using new link names are activated.
If there's a filtering rule that applies to interface bge0, and someone renames
bge0 to net0, then the rule no longer matches packets received on the link
formally known as bge0.

Also IPFilter has been designed to allow users to specify rules with interface
names that do not exist at the time they are loaded, and for those interface
names to be resolved at the time at which they're added to the system. Thus,
the mapping from the linkname to the linkid needs to happen in the kernel.
Changing IPFilter to use linkid instead of link name will not work.

Protocol & Hook registration
============================
Protocol registration
---------------------
With IP layer netinfo today we have 3 protocols, IPv4, IPv6 and ARP. For MAC
layer, each of the MAC plugin type is treated as a different protocol, so 
we'll have ethernet, wifi and ib. These protocols will be registered by
using net_protocol_register() when the corresponding MAC plugin gets loaded.

Hook registration
-----------------
IPFilter will register hooks for MAC layer protocols in the following cases:

when the first ethernet filtering rule is added
- register the ethernet hook

when the first "layer2" IP filtering/NAT rule is added
- register the ethernet, wifi and ib hooks

when it receives a notification indicating that a protocol is registered
- register the hook if there are rules for that corresponding protocol.

Since layer 2 filtering functionality is enabled automatically when the 
first layer 2 rule is added, the corresponding hook needs to be registered
then so packets can be passed to IPFilter from the hook framework.

It is possible that a rule for a layer 2 protocol is added before the
corresponding protocol is registered. Suppose user has added a layer 2 IP
filtering rule on a system that only has ethernet cards, then he plugs a
wifi card into the system and sets it up, in this case when the wifi MAC
plugin is loaded, the protocol will be registered, and IPFilter will be
notified via the callback notification mechanism provided by the PFHooks
API project, and it will register the hook for that protocol so it can
receive and match wifi packets.

| Dynamic data path modification
| ------------------------------
| To make sure layer 2 filtering has no performance impact when disabled,
| instead of inserting hooks check into the fast path code, we make use of
| the function pointer driven approach provided by Crossbow where possible.
| On RX side layer 2 filter implements its own receive function, and will
| replace the default function with its own one when l2 filtering is enabled.
| The l2 filter specfic receive function does layer 2 firewall processing
| before calling the original receive function. So when l2 filter is disabled
| there's zero additional processing on the RX path. On TX side layer 2 filter
| will force packets off the fast path when filtering is enabled, and add
| the hooks check into the non fast path to avoid impacting performance.
| In both cases the data path will be modified dynamically when the filtering
| is enabled/disabled, and this is done on a per MAC client basis.
| 
| To do this a function will need to be called when the first hook is registered
| on a specific hook event, and when the last hook is unregistered from the
| event, to do the the necessary data path setup. The hook_event_t strcture is
| changed to accomodate this so that hook providers, MAC plugins in this case,
| could specify their own callbacks which will be called from hooks_register/
| hook_unregister().
| 
| typedef struct hook_event_s {
|          int             he_version;
|          char            *he_name;       /* name of this hook list */
|          int             he_flags;       /* 1 = multiple entries allowed */
|          boolean_t       he_interested;  /* true if callback exist */
| +        void            (*he_enable_cb)(hook_event_token_t, hook_event_t *,
| +                            void *);
| +        void            (*he_disable_cb)(hook_event_token_t, hook_event_t *,
| +                            void *);
| +        void            *he_arg_cb;
| } hook_event_t;
| 
| he_arg_cb points at a mactype_t structure, to identify which MAC plugin the
| hook is registered on, as l2 filtering is enabled/disabled per MAC plugin.
| The two callback functions, pointed by he_enable_cb and he_disable_cb, will
| walk through the MAC clients in the system, and does the necessary data path
| setup/cleanup for the corresponding MAC clients, which are primary MAC clients
| on top of data links of the specific MAC plugin.
|
| Relative Hooks ordering
| =======================
| Order with Bridging
| -------------------
| L2 filtering is done on a per MAC client basis. When the users specify "net0",
| this refers to the traffic going through the primary MAC client of net0, for
| example IP on top of that data link. This is different from all traffic going
| through the physical MAC instance which is shared by multiple MAC clients.
| And L2 hooks intercept traffic both from/to the wire, and those occur between
| multiple MAC clients defined on top of the same data link.
| 
| With regard to bridging, L2 filter works on top of the bridge, instead of
| underneath it, as the filtering is based on MAC clients instead of the MAC
| instance that the bridge uses. This means in certain cases L2 hooks is not
| able to see the actual interface used for transmit or receive, but only the
| interface that the network layer believes it's using, as when IP sends a
| packet on one interface, the bridge may end up transmitting that packet
| on another interface - if that's the interface on which the destination
| exists or if the destination is unknown. This is by design as l2 filtering
| aims more on controling what packets a VM can send to the wire via the data
| link it is using, insted of which physical link the packets actually get sent
| out from.
| 
| Order with bandwidth limit
| --------------------------
| L2 filter sits underneath bandwidth shaping by Crossbow. On RX side, the
| filtering is done before the bandwidth limit is applied; on TX side, it
| is applied after the bandwidth limit. 

IPFilter changes
================
Users can use ipf(1M) to add ethernet filtering rules in addition to IP 
filtering rules, these ethernet filtering rules are marked with "family ether".
They can also add IP Filtering/NAT rules and mark them with "layer2" keyword
so these rules will be processed in MAC layer instead of IP layer. Unlike IPv6,
no special command line switch is required to load these rules.

The "layer2" IP filtering/NAT rules go to existing ipf.conf, ipf6.conf and
ipnat.conf, respectively. The "family ether" rules go to a new configuration
file ipf-ether.conf.

The layer 2 filtering functionality will be enabled automatically when the
first ethernet rule or "layer2" IPFilter rule is added, and disabled when
the last such rule is removed. This functionality is only available in
global zone.

Also, ipmon has been updated to print out log records with ethernet
information but the output of this command is volatile.

Rule processing 
---------------
Currently processing order in IPFilter is:

[INPUT] -> IP NAT -> IP firewall -> { IP }  -> IP firewall -> IP NAT -> [OUTPUT]

With layer 2 filtering the processing order would become:

| [INPUT] -> L2 firewall -> "layer2" IP NAT -> "layer2" IP firewall ->
| ... -> IP NAT -> IP firewall -> { IP }  -> IP firewall -> IP NAT -> ...
| -> "layer2" IP firewall -> "layer2" IP NAT -> L2 firewall -> [OUTPUT]

Input processing
~~~~~~~~~~~~~~~~
Take input processing for an IP packet for example:

- MAC level filtering rules are processed first. These rules match on MAC
headers to determine if a packet should be passed or blocked. Administrators
use these rules to match with MAC addresses, MAC type, VLAN ID, .etc.

- L2filter jump over the MAC header, determine if this is an IP packet, and
do some sanity checking before passing it up to "layer2" IP rules for further
processing.

- Then "layer2" IP NAT rules are processed. Like IP layer NAT rules, these
rules do NAT for IP packets, but it is done at MAC layer instead of IP layer.

- Then "layer2" IP Filtering rules are processed. These rules provide IP 
Filtering at MAC layer.

- L2filter finishes processing and the packet is delivered up in the stack.
When the packet reaches IP, IP layer filtering/NAT processing is invoked,
and it works just as it does today.

Changes to output
-----------------
With layer 2 filtering, each type of rules have its own distinct orders,
the output of ipfstat/ipnat has been modified so that the rules are shown
in a manner to let the users better understand the processing orders. 
The change only applies to global zone, output in non-global zones remain
unchanged.

Example
~~~~~~~

# ipfstat -io
Ethernet rules:
empty list for ipfilter(out)
pass in family ether all
pass in family ether from 1:2:3:4:5:6 to any

layer 2 IP rules:
empty list for ipfilter(out)
pass in proto icmp from 1.1.1.1 to 2.2.2.2 layer2
block in proto tcp from 3.3.3.3 to 4.4.4.4 layer2

IP rules:
pass in all
pass out all

# ipnat -l
List of layer 2 active MAP/Redirect filters:
map bge1 from 2.3.4.5/32 to 6.7.8.9/32 -> 1.1.2.2/32 layer2

List of active MAP/Redirect filters:

List of active sessions:

Examples
--------
Prevent MAC address spoofing
~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Suppose we have a domU with a interface vnic0, we may want to ensure:
- packets from this domU can have use its own source MAC address, preventing
this domU from pretending someone else
- packet from this source MAC address can only come from this domU, preventing
others from pretending this domU

say vnic0 has MAC address 11:22:33:44:55:66, the rules would be something like:

block out family ether from 11:22:33:44:55:66 to any
block out on vnic0 family ether from any to any
pass out on vnic0 family ether from 11:22:33:44:55:66 to any

Prevent IP address spoofing
~~~~~~~~~~~~~~~~~~~~~~~~~~~

Suppose we'd want to prevent a domU from using others' IP addresses, we can
probably go with:

block out on vnic0 from any to any layer2
pass out on vnic0 from 1.1.1.1 to any layer2

while 1.1.1.1 is the assigned IP address on vnic0

VLAN packets filtering
~~~~~~~~~~~~~~~~~~~~~~

Say we'd want to block some IP traffic, below are examples on how it is done
with regard to vlan:

- block all IP traffic regardless of VLAN

block in family ether type 0x800 

- block all IP traffic belonging to VLAN:

block in family ether type 0x800 with vlan

- block all IP traffic NOT belonging to VLAN:

block in family ether type 0x800 with not vlan 

- block all IP traffic for a specific VLAN (e.g. 100)

block in family ether type 0x800 vlan 100

Ioctl compatibility
-------------------
ABI compatibility with the old structure definitions is preserved by this case.

IPFILTER_VERSION (see ipnat(7i)) is used to keep track of user application's
version thus the old binaries can still work after this change. The kernel code
would handle the ioctl input/output based on the version number to make it a
compatible change. There's no change required for user applications using the
interfaces.

The related data structures natlookup_t and nat_t remain the same, and ioctls
SIOCGNATL/SIOCSTPUT will work correctly. User can set a flag, IPN_LAYER2,
in natlookup_t and nat_t, respectively, to indicate it is looking up/inserting
a layer 2 NAT session, or a layer 3 one. For compatibilities, by default the
flag is not set, which indicates a layer 3 session.


Interfaces
==========
+----------------------------------------+-------------------+
| Interface                              |  Classification   |
+----------------------------------------+-------------------+
| NE_NAME_CHANGE                         |     Committed     |
| NHF_ETHER                              |     Committed     |
| NHF_WIFI                               |     Committed     |
| NHF_IB                                 |     Committed     |
| | <sys/hook.h>                         |     Committed     |		
| | <sys/hook_event.h>                   |     Committed     |		
| | <sys/neti.h>                         |     Committed     |		
+----------------------------------------+-------------------+
| "ipfilter_hook_eth_in"                 |    Uncommitted    |
| "ipfilter_hook_eth_out"                |    Uncommitted    |
| "ipfilter_hook_wifi_in"                |    Uncommitted    |
| "ipfilter_hook_wifi_out"               |    Uncommitted    |
| "ipfilter_hook_ib_in"                  |    Uncommitted    |
| "ipfilter_hook_ib_out"                 |    Uncommitted    |
+----------------------------------------+-------------------+
| "family ether"                         |      Committed    |
| "layer2"                               | Obsolete Volatile |
+----------------------------------------+-------------------+
| IPN_LAYER2                             |      Volatile     |
| /usr/include/netinet/ip_fil.h          |    Uncommitted    |
| /usr/include/netinet/ip_nat.h          |    Uncommitted    |
+----------------------------------------+-------------------+


From Sebastien.Roy@Sun.COM Tue Jan 20 10:17:00 2009
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n0KIGx91019902
	for <psarc-ext@sac.sfbay.sun.com>; Tue, 20 Jan 2009 10:16:59 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id n0KIGwbT041987
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Tue, 20 Jan 2009 11:16:59 -0700 (MST)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KDS00D0L84A6H00@nwk-avmta-1.sfbay.Sun.COM> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Tue, 20 Jan 2009 10:16:58 -0800 (PST)
Received: from brmea-mail-4.sun.com ([192.18.98.36])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KDS000TQ8465K60@nwk-avmta-1.sfbay.Sun.COM> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Tue,
 20 Jan 2009 10:16:55 -0800 (PST)
Received: from fe-amer-10.sun.com ([192.18.109.80])
	by brmea-mail-4.sun.com (8.13.6+Sun/8.12.9) with ESMTP id n0KIGsWO012263	for
 <PSARC-ext@sun.com>; Tue, 20 Jan 2009 18:16:54 +0000 (GMT)
Received: from conversion-daemon.mail-amer.sun.com by mail-amer.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KDS00B0176J0Z00@mail-amer.sun.com>
 (original mail from Sebastien.Roy@Sun.COM)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Tue,
 20 Jan 2009 11:16:54 -0700 (MST)
Received: from [129.148.174.103] by mail-amer.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KDS00B6D846GZ00@mail-amer.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Tue, 20 Jan 2009 11:16:54 -0700 (MST)
Date: Tue, 20 Jan 2009 13:16:44 -0500
From: Sebastien Roy <Sebastien.Roy@Sun.COM>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <4975EE39.9080906@Sun.COM>
Sender: Sebastien.Roy@Sun.COM
To: Darren Reed <Darren.Reed@Sun.COM>
Cc: PSARC-EXT <PSARC-ext@Sun.COM>, Zhijun Fu <Zhijun.Fu@Sun.COM>
Message-id: <1232475404.9787.9.camel@strat>
Organization: Sun Microsystems
MIME-version: 1.0
X-Mailer: Evolution 2.24.2
Content-type: text/plain; charset=UTF-8
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM> <4975EE39.9080906@Sun.COM>
Status: RO
Content-Length: 565

On Tue, 2009-01-20 at 16:31 +0100, Darren Reed wrote:
> | MAC client index
> | ----------------
> | L2 filtering is based on MAC client which is introduced by Crossbow project,
> | and the filtering is done on a per MAC client basis. When users specify a
> | link name "net0", this corresponds to the traffic going through the primary
> | MAC client of net0, e.g. IP on top of that data link. 

How does this work with bridging (PSARC 2008/055)?  When the bridge
forwards packets between two MAC providers, there's presumably no MAC
client involved at all.

-Seb



From carlsonj@phorcys.east.sun.com Tue Jan 20 10:26:14 2009
Received: from sunmail2sca.sfbay.sun.com (sunmail2sca.SFBay.Sun.COM [129.145.155.234])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n0KIQENO020080
	for <psarc-ext@sac.sfbay.sun.com>; Tue, 20 Jan 2009 10:26:14 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail2sca.sfbay.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id n0KIQ9XM008426;
	Tue, 20 Jan 2009 10:26:11 -0800 (PST)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KDS0043N8JM7400@nwk-avmta-2.sfbay.sun.com>; Tue,
 20 Jan 2009 10:26:10 -0800 (PST)
Received: from phorcys.east.sun.com ([129.148.174.143])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KDS00LKW8JL0G90@nwk-avmta-2.sfbay.sun.com>; Tue,
 20 Jan 2009 10:26:10 -0800 (PST)
Received: from phorcys.east.sun.com (localhost [127.0.0.1])
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3) with ESMTP id n0KIQ95k022396; Tue,
 20 Jan 2009 13:26:09 -0500 (EST)
Received: (from carlsonj@localhost)
	by phorcys.east.sun.com (8.14.3+Sun/8.14.3/Submit) id n0KIQ9vo022393; Tue,
 20 Jan 2009 13:26:09 -0500 (EST)
Date: Tue, 20 Jan 2009 13:26:09 -0500
From: James Carlson <james.d.carlson@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <1232475404.9787.9.camel@strat>
To: Sebastien Roy <Sebastien.Roy@sun.com>
Cc: Darren Reed <Darren.Reed@sun.com>, PSARC-EXT <PSARC-ext@sun.com>,
        Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <18806.5953.414729.404636@gargle.gargle.HOWL>
MIME-version: 1.0
X-Mailer: VM 7.01 under Emacs 21.3.1
Content-type: text/plain; charset=us-ascii
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM> <4975EE39.9080906@Sun.COM>
 <1232475404.9787.9.camel@strat>
Status: RO
Content-Length: 981

Sebastien Roy writes:
> On Tue, 2009-01-20 at 16:31 +0100, Darren Reed wrote:
> > | MAC client index
> > | ----------------
> > | L2 filtering is based on MAC client which is introduced by Crossbow project,
> > | and the filtering is done on a per MAC client basis. When users specify a
> > | link name "net0", this corresponds to the traffic going through the primary
> > | MAC client of net0, e.g. IP on top of that data link. 
> 
> How does this work with bridging (PSARC 2008/055)?  When the bridge
> forwards packets between two MAC providers, there's presumably no MAC
> client involved at all.

That's correct.  Filtering at this level won't catch bridge-forwarded
packets.

I think the answer is that we'll need proper hooks in the forwarding
path.

-- 
James Carlson, Solaris Networking              <james.d.carlson@sun.com>
Sun Microsystems / 35 Network Drive        71.232W   Vox +1 781 442 2084
MS UBUR02-212 / Burlington MA 01803-2757   42.496N   Fax +1 781 442 1677

From Darren.Reed@sun.com Wed Jan 21 04:45:15 2009
Received: from sunmail5.uk.sun.com (sunmail5.UK.Sun.COM [129.156.85.165])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n0LCjEwt006928
	for <psarc-ext@sac.sfbay.sun.com>; Wed, 21 Jan 2009 04:45:15 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail5.uk.sun.com (8.13.8+Sun/8.13.8/ENSMAIL,v2.2) with ESMTP id n0LCj81x005286
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Wed, 21 Jan 2009 12:45:13 GMT
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KDT0050ZNFBIB00@nwk-avmta-2.sfbay.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Wed, 21 Jan 2009 04:45:11 -0800 (PST)
Received: from gmp-eb-inf-2.sun.com ([192.18.6.24])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KDT0050SNFAH800@nwk-avmta-2.sfbay.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Wed,
 21 Jan 2009 04:45:11 -0800 (PST)
Received: from fe-emea-09.sun.com (gmp-eb-lb-2-fe3.eu.sun.com [192.18.6.12])
	by gmp-eb-inf-2.sun.com (8.13.7+Sun/8.12.9) with ESMTP id n0LCj9xW014694	for
 <PSARC-ext@sun.com>; Wed, 21 Jan 2009 12:45:09 +0000 (GMT)
Received: from conversion-daemon.fe-emea-09.sun.com by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KDT00201MOA9G00@fe-emea-09.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Wed,
 21 Jan 2009 12:45:09 +0000 (GMT)
Received: from [129.157.18.18] by fe-emea-09.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KDT004H8NEN7HD0@fe-emea-09.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Wed, 21 Jan 2009 12:44:47 +0000 (GMT)
Date: Wed, 21 Jan 2009 13:44:44 +0100
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <4975EE39.9080906@Sun.COM>
Sender: Darren.Reed@sun.com
To: Zhijun Fu <Zhijun.Fu@sun.com>
Cc: PSARC-EXT <PSARC-ext@sun.com>
Message-id: <497718BC.8060206@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM> <4975EE39.9080906@Sun.COM>
User-Agent: Thunderbird 2.0.0.19 (Windows/20081209)
Status: RO
Content-Length: 1470

...
> | Relative Hooks ordering
> | =======================
> | Order with Bridging
> | -------------------
...
> With regard to bridging, L2 filter works on top of the bridge, instead of
> underneath it, as the filtering is based on MAC clients instead of the 
> MAC
> instance that the bridge uses. This means in certain cases L2 hooks is 
> not
> able to see the actual interface used for transmit or receive, but 
> only the
> interface that the network layer believes it's using, as when IP sends a
> packet on one interface, the bridge may end up transmitting that packet
> on another interface - if that's the interface on which the destination
> exists or if the destination is unknown. This is by design as l2 
> filtering
> aims more on controling what packets a VM can send to the wire via the 
> data
> link it is using, insted of which physical link the packets actually 
> get sent
> out from.

Let me spell out the issue here: without being able to reliably
associate packets presented by the layer 2 packet events, it is
not possible to use them with bridging to implement a security
device, be it with IPFilter or something else.

If the packets published by the layer 2 hooks NH_PHYSICAL_IN
and NH_PHYSICAL_OUT do not correspond 100% of the time to the
actual physical interfaces that the packets are sent out on
and received on, what events would? And therefore what should
the packet events being delivered by this project really be
called?

Darren


From Darren.Reed@sun.com Thu Jan 22 02:51:01 2009
Received: from sunmail4.singapore.sun.com (sunmail4.Singapore.Sun.COM [129.158.71.19])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n0MAp0gE013759
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 22 Jan 2009 02:51:01 -0800 (PST)
Received: from nwk-avmta-2.sfbay.sun.com (nwk-avmta-2.SFBay.Sun.COM [129.145.155.6])
	by sunmail4.singapore.sun.com (8.13.4+Sun/8.13.3/ENSMAIL,v2.2) with ESMTP id n0MAoXu0008713
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Thu, 22 Jan 2009 18:50:59 +0800 (SGT)
Received: from pmxchannel-daemon.nwk-avmta-2.sfbay.sun.com by
 nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KDV00G0LCSXZG00@nwk-avmta-2.sfbay.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Thu, 22 Jan 2009 02:50:57 -0800 (PST)
Received: from gmp-eb-inf-2.sun.com ([192.18.6.24])
 by nwk-avmta-2.sfbay.sun.com
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KDV00A5VCSW5T80@nwk-avmta-2.sfbay.sun.com> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Thu,
 22 Jan 2009 02:50:56 -0800 (PST)
Received: from fe-emea-10.sun.com (gmp-eb-lb-2-fe2.eu.sun.com [192.18.6.11])
	by gmp-eb-inf-2.sun.com (8.13.7+Sun/8.12.9) with ESMTP id n0MAotwV021348	for
 <PSARC-ext@sun.com>; Thu, 22 Jan 2009 10:50:55 +0000 (GMT)
Received: from conversion-daemon.fe-emea-10.sun.com by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KDV00J01BKADJ00@fe-emea-10.sun.com>
 (original mail from Darren.Reed@Sun.COM)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Thu,
 22 Jan 2009 10:50:55 +0000 (GMT)
Received: from [129.157.18.18] by fe-emea-10.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KDV0032YCST7U00@fe-emea-10.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Thu, 22 Jan 2009 10:50:53 +0000 (GMT)
Date: Thu, 22 Jan 2009 11:50:50 +0100
From: Darren Reed <Darren.Reed@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <497718BC.8060206@Sun.COM>
Sender: Darren.Reed@sun.com
To: PSARC-EXT <PSARC-ext@sun.com>
Cc: Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <49784F8A.3010800@Sun.COM>
MIME-version: 1.0
Content-type: text/plain; format=flowed; charset=ISO-8859-1
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM> <4975EE39.9080906@Sun.COM>
 <497718BC.8060206@Sun.COM>
User-Agent: Thunderbird 2.0.0.19 (Windows/20081209)
Status: RO
Content-Length: 1921

I'm putting this fast track back into "waiting need spec" as the
submitter has taken this conversation offline.

If there any members reading this, I'd like them to consider
derailing this fast track as, in my opinion, it has become less
than obvious (this is the 3rd time to "waiting need spec") due
to the requirements and interactions being less than trivial to
sort out.

Darren Reed wrote:
> ...
>> | Relative Hooks ordering
>> | =======================
>> | Order with Bridging
>> | -------------------
> ...
>> With regard to bridging, L2 filter works on top of the bridge, 
>> instead of
>> underneath it, as the filtering is based on MAC clients instead of 
>> the MAC
>> instance that the bridge uses. This means in certain cases L2 hooks 
>> is not
>> able to see the actual interface used for transmit or receive, but 
>> only the
>> interface that the network layer believes it's using, as when IP sends a
>> packet on one interface, the bridge may end up transmitting that packet
>> on another interface - if that's the interface on which the destination
>> exists or if the destination is unknown. This is by design as l2 
>> filtering
>> aims more on controling what packets a VM can send to the wire via 
>> the data
>> link it is using, insted of which physical link the packets actually 
>> get sent
>> out from.
>
> Let me spell out the issue here: without being able to reliably
> associate packets presented by the layer 2 packet events, it is
> not possible to use them with bridging to implement a security
> device, be it with IPFilter or something else.
>
> If the packets published by the layer 2 hooks NH_PHYSICAL_IN
> and NH_PHYSICAL_OUT do not correspond 100% of the time to the
> actual physical interfaces that the packets are sent out on
> and received on, what events would? And therefore what should
> the packet events being delivered by this project really be
> called?
>
> Darren
>


From Sebastien.Roy@sun.com Thu Jan 22 09:09:07 2009
Received: from newsunmail1brm.central.sun.com (newsunmail1brm.Central.Sun.COM [129.147.62.245])
	by sac.sfbay.sun.com (8.13.8+Sun/8.13.8) with ESMTP id n0MH96PE025363
	for <psarc-ext@sac.sfbay.sun.com>; Thu, 22 Jan 2009 09:09:07 -0800 (PST)
Received: from nwk-avmta-1.SFBay.Sun.COM (nwk-avmta-1.SFBay.Sun.COM [129.146.11.74])
	by newsunmail1brm.central.sun.com (8.13.7+Sun/8.13.7/ENSMAIL,v2.2) with ESMTP id n0MH93ZC015275
	for <@sunmail2sca.sfbay.sun.com:PSARC-ext@sun.com>; Thu, 22 Jan 2009 10:09:06 -0700 (MST)
Received: from pmxchannel-daemon.nwk-avmta-1.sfbay.Sun.COM by
 nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 id <0KDV00A2DUB4PG00@nwk-avmta-1.sfbay.Sun.COM> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Thu, 22 Jan 2009 09:09:04 -0800 (PST)
Received: from brmea-mail-2.sun.com ([192.18.98.43])
 by nwk-avmta-1.sfbay.Sun.COM
 (Sun Java System Messaging Server 6.2-3.04 (built Jul 15 2005))
 with ESMTP id <0KDV00AJCUB31Y10@nwk-avmta-1.sfbay.Sun.COM> for
 PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Thu,
 22 Jan 2009 09:09:03 -0800 (PST)
Received: from fe-amer-09.sun.com ([192.18.109.79])
	by brmea-mail-2.sun.com (8.13.6+Sun/8.12.9) with ESMTP id n0MH92DL018461	for
 <PSARC-ext@sun.com>; Thu, 22 Jan 2009 17:09:02 +0000 (GMT)
Received: from conversion-daemon.mail-amer.sun.com by mail-amer.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 id <0KDV00D01T8A8T00@mail-amer.sun.com>
 (original mail from Sebastien.Roy@Sun.COM)
 for PSARC-ext@sun.com (ORCPT PSARC-ext@sun.com); Thu,
 22 Jan 2009 10:09:02 -0700 (MST)
Received: from [129.148.174.103] by mail-amer.sun.com
 (Sun Java System Messaging Server 6.2-8.04 (built Feb 28 2007))
 with ESMTPSA id <0KDV00G1DUANPA50@mail-amer.sun.com> for PSARC-ext@sun.com
 (ORCPT PSARC-ext@sun.com); Thu, 22 Jan 2009 10:08:48 -0700 (MST)
Date: Thu, 22 Jan 2009 12:08:37 -0500
From: Sebastien Roy <Sebastien.Roy@sun.com>
Subject: Re: PSARC/2008/249 Packet interception for the MAC layer
In-reply-to: <49784F8A.3010800@Sun.COM>
Sender: Sebastien.Roy@sun.com
To: Darren Reed <Darren.Reed@sun.com>
Cc: PSARC-EXT <PSARC-ext@sun.com>, Zhijun Fu <Zhijun.Fu@sun.com>
Message-id: <1232644117.11415.136.camel@strat>
Organization: Sun Microsystems
MIME-version: 1.0
X-Mailer: Evolution 2.24.2
Content-type: text/plain; charset=UTF-8
Content-transfer-encoding: 7BIT
X-PMX-Version: 5.4.1.325704
References: <47FC5A55.2050405@Sun.COM> <4975EE39.9080906@Sun.COM>
 <497718BC.8060206@Sun.COM> <49784F8A.3010800@Sun.COM>
Status: RO
Content-Length: 1545

On Thu, 2009-01-22 at 11:50 +0100, Darren Reed wrote:
> I'm putting this fast track back into "waiting need spec" as the
> submitter has taken this conversation offline.

Okay, thanks.

> If there any members reading this, I'd like them to consider
> derailing this fast track as, in my opinion, it has become less
> than obvious (this is the 3rd time to "waiting need spec") due
> to the requirements and interactions being less than trivial to
> sort out.

There's one outstanding issue being discussed offline, and that is
related to this case's lack of layer-2 filtering on a MAC provider
basis.  This results in this project not being able to intercept some
packets (those forwarded between physical links by a bridge), and
doesn't allow the administrator the flexibility to easily express rules
that apply to all MAC clients for a given MAC provider (i.e., filter all
packets coming in and out of "bge0").  My assertion is that this case is
thus incomplete as was proposed.

These issues can likely be resolved, and the resolution can be provided
in the form of a fast-track.  The details of what the resolution looks
like need to be designed outside the context of ARC review anyway, so
I'm not convinced that a full review will help very much with that.  If
the project team feels that they'll save some time by having an
in-meeting case review, then by all means, you're free to request a
review slot and get a full review.

Otherwise, it seems fine to me to work out the details of the missing
pieces and submit new materials.

-Seb



